Life sciences · Preprint
arXiv · September 8, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint reports a computational study using a frozen-rater design applied to 124,615 ICLR peer reviews spanning 2018–2025. The authors find that human reviewers' reward for non-domain lexical complexity declined from +0.142 to −0.015 over the period, while a machine model evaluated all submissions identically in one window, yielding a three-way difference-in-differences estimate of −0.0100 (q=0.013). The result suggests reviewers discounted lexical elaboration as its production cost fell, but the study does not directly measure reviewer preferences or demonstrate causation.
Computational observational study with frozen-rater comparison design. 32,638 unique submissions to ICLR across 2018–2025 with 124,615 human peer reviews; submissions span the full range of papers accepted to and rejected by the conference.. Intervention: Frozen-rater machine model (one model family, one prompt, evaluated all submissions in a single February–April 2025 window) used to generate counterfactual reviews with no temporal variation in model or evaluation logic.. Compared with: Human peer reviews from the same submissions, evaluated separately for each year 2018–2025, to detect year-to-year drift in the relationship between lexical complexity and review scores.. n = 32,638. ICLR conference (venue not geographically specified, but ICLR is an international venue with reviews from global reviewers)..
Human coefficient on non-domain lexical complexity fell from +0.142 to −0.015 across 2018–2025 Frozen rater coefficient remained stable at +0.080 to +0.082 across the same period Three-way difference-in-differences: −0.0100 (q=0.013), indicating preference drift net of submission composition change
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
A novel observational study using computational methods to infer reviewer preference drift from 8 years of peer review data, with clever design but no direct measurement of causal mechanism and applicability limited to one venue and evaluation model.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Large language models have collapsed the cost of producing lexically elaborate prose, and whether peer reviewers still reward it is a question about the evaluator, not about the text. When the association between a writing cue and review scores moves across years, the reviewers may have changed, the submissions may have changed, or both, and a regression of scores on text cannot say which. We separate the two with a frozen rater: 81,850 machine reviews of ICLR submissions from 2018 to 2025, all generated in one February-April 2025 window with one model family and one prompt, so that its year-to-year coefficients track submission composition alone and the human-minus-frozen trend difference identifies reviewer preference drift. On 32,638 submissions with 124,615 human reviews, the human coefficient on non-domain lexical complexity falls from +0.142 to -0.015 while the frozen rater moves from +0.080 to +0.082; the three-way difference-in-differences is -0.0100 (q=0.013), and forty random-wordlist placebos through the same specification centre on zero. Humans still reward sentence-length variability, which the frozen rater never registers, while the frozen rater still pays for lexical complexity at its earlier rate. Every claim is held to a double gate of false-discovery control and interval exclusion, and the findings that failed adversarial re-testing are reported. Reviewers discounted a cue whose production cost collapsed, as models of manipulable signals prescribe; an LLM judge calibrated to historical human preferences inherits the earlier schedule and drifts out of alignment while its agreement with humans on totals stays ordinary.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.