Life sciences · Preprint
arXiv · September 9, 2026
Raises a question worth testing. It does not answer one.
This preprint proposes a theoretical framework (joint-Gaussian model of verifier-gold correlation) and empirical approach ('proof-carrying cognition' with reality-settled rewards) to reduce reward hacking in language-model reasoning. Experiments show that unsound verifiers degrade under optimization pressure in program synthesis, while reality-anchored settlement and stronger judges are more robust, with potential ~10-fold label efficiency gain. However, the work remains unrefereed, is confined to formal domains with executable ground truth, lacks clear applicability to open-ended reasoning, and does not report absolute sample sizes or generalization evidence.
Mixed methods: theoretical joint-Gaussian model, controlled program-synthesis experiments with executable ground truth, and uncontrolled LLM judge robustness assessments. Program-synthesis tasks with executable ground truth; language models with weak and stronger LLM judges; systems trained with GRPO. No human subjects or clinical population.. Intervention: Reality-anchored settlement of reasoning rewards; refit reward model on 10% settlement stream; proper scoring rules; self-built world model trained on held-out reality.. Compared with: Frozen verifier; unsound verifier; weak LLM judge; random labeling; selection alone without reward model refit..
Unsound verifier soundness degrades from 0.94 to 0.32 at N=4096 under optimization; sound verifier improves monotonically. Reality-anchored settlement reduces hacking gap from ~0.27 to ~0 under adversarial pressure. Soundness scales log-linearly with settled labels; on-policy settlement ~10x more label-efficient than random labeling.
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This preprint presents theoretical modeling, controlled demonstrations in formal domains (program synthesis), and a proposed paradigm without peer review or evidence of real-world clinical or deployed reasoning system validation.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Frontier gains in language-model reasoning come from reinforcement learning on reasoning traces and are concentrated in domains with a cheap, sound verifier. We argue the field's binding constraint is the verification gap: no scalable, incorruptible reward for reasoning outside formal domains. We make four contributions. (1) Theory: in a joint-Gaussian model of best-of-N selection, verifier-gold correlation rho is the exact exchange rate between test-time compute and capability, and an unsound verifier pays a polynomial penalty N^(1/rho^2); a margin-free copula form predicts realized soundness of real LLM judges to 4% median error. (2) Demonstration: in program-synthesis testbeds with executable ground truth, including a pre-registered scaled replication, unsound verifiers lose Soundness-under-Pressure as optimization grows (0.94 to 0.32 at N=4096) while a sound verifier improves monotonically; reality-anchored settlement beats a frozen verifier under i.i.d. and adversarial pressure, driving the hacking gap from ~0.27 to ~0; soundness scales log-linearly with settled labels, with on-policy settlement ~10x more label-efficient than random labeling. With real LLM judges and unit-test execution as gold, a weak judge loses soundness under best-of-N (p<0.001), a stronger judge is more robust, and selection alone manufactures +0.53 hacking gaps from honest samples. Under real GRPO training, a frozen reward model traces the full overoptimization curve (executed reward collapses 90%) while the same model refit on a 10% settlement stream preserves 6x the executed reward. (3) Paradigm: proof-carrying cognition, where reasoning steps are typed probabilistic claims priced by a self-built world model trained only on held-out reality and settled by proper scoring rules. (4) Benchmark: we specify Soundness-under-Pressure as the headline metric for a reality-settled reasoning benchmark.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.