Zenodo
Dataset
DOI 10.5281/zenodo.22099066 (opens in a new tab)
Abstract
A reproducible measurement instrument that captures the human-centred quality of language-model decision support and the system cost of producing it over the same generation, in a single pass: an Empathy–Trust–Accountability rubric scored by a fixed automated judge (gpt-4o-mini), alongside per-inference latency, output throughput and resident memory. The deposit contains the measurement harness; the judge prompts and the full ETA rubric verbatim; twelve high-stakes health-decision scenarios parameterized with published NFHS-5 (2019–21) all-India survey indicators, with sources; the 144-row evaluation dataset (12 scenarios × 4 quantized Llama-3.2 variants × 3 runs, none excluded) from which everyreported number derives, and the analysis script that produces every statistic, table, and figure in the associated article. Reproducibility is bounded, and the bounds are stated in the README: the analysis is deterministic from the released dataset, but the dataset is not deterministic from the models — re-running the benchmark re-runs live generation and live judging, neither of which is seeded. The energy column is present and zero-filled. Model tags are not pinned to manifest digests. The harness is pinned by commit rather than by a release tag. Code is licensed GPL-2.0-or-later; data are licensed CC BY 4.0.