Life sciences · Preprint
arXiv · September 3, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preregistered audit demonstrates that language models used as measurement instruments on shared serving endpoints do not achieve required reliability (Spearman 0.400 vs. required 0.90 for same-window repeats; 0.78 vs. required 0.99 for byte-identical next-day replays). The instability arises from label-to-meaning mapping bias, candidate ranking gaps below instrument noise floor, and byte-identical inputs returning different rankings—none repaired by metric substitution, sampling, or waiting. The finding is methodologically rigorous but unrefereed and concerns instrumental validity in AI evaluation, not clinical or human health outcomes.
Preregistered reliability audit with dual-campaign protocol and fixed thresholds. Language model endpoints on shared serving infrastructure; four providers tested across multiple conditions. Intervention: Repeated identical requests (same-window and byte-identical next-day) sent to same model name on shared endpoints; systematic follow-up conditions including provider switching, self-hosting, and constructed errors. Compared with: Preregistered reliability thresholds (Spearman 0.90 for same-window repeats; 0.99 for byte-identical replays). n = 52,988. Not stated.
Same-window repeat rankings agreed at Spearman 0.400 against required 0.90 across 52,988 audited request attempts Byte-identical next-day replays agreed at Spearman 0.78 against required 0.99 Four providers tested showed medians ranging 0.74 to 0.88 on tested grid, with none predicted by exposed metadata fields
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
A preregistered audit of measurement reliability in LLM-based evaluation systems, demonstrating systematic failure to meet validity thresholds; methodologically sound but non-peer-reviewed, addressing instrumental rather than clinical validity.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurement instrument, resting on one rarely stated assumption: the same request, sent to the same model name, reads the same tomorrow. We audited that assumption in two preregistered campaigns with every threshold fixed in advance; neither got past validating its instrument. Across 52,988 audited request attempts, same-window repeat rankings agreed at Spearman 0.400 against a required 0.90, and byte-identical next-day replays agreed at 0.78 against a required 0.99, each time with the execution record at ceiling. Three mechanisms explain the gap: a label-to-meaning mapping that biased readouts as strongly as the signal; candidate gaps seven orders of magnitude below the instrument's own noise floor; and byte-identical inputs returning different rankings, a noise that exact-permutation readouts compound. Neither metric substitution nor sampling repaired it on the tested grid. Preregistered follow-ups bound the problem: waiting did not help on the days sampled (0.805 versus 0.800, replicated over five further days); switching providers did not help (four providers share the floor, medians 0.74 to 0.88, predicted by none of the metadata fields they expose); self-hosting on batch-invariant kernels helped only while the server was quiet; and on constructed errors with known gaps, the readout's separation tracks error type, not size. We distill the evidence into a three-level snapshot-identity ladder, eight design rules, and a reporting checklist; a pilot at roughly 2% of the study's call volume would have exposed both unreachable gates in advance. All results concern externally measured behaviour on shared serving infrastructure. On a shared endpoint, a model name is not a frozen instrument; a preregistered evaluation must measure its instrument before freezing any gate on it.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.