Life sciences · Preprint
arXiv · August 14, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint identifies a paradox in Power Sampling—a technique that can shift probability mass toward correct language model trajectories yet degrade downstream reasoning accuracy—and proposes a deformation-controlled alternative. The work is computational and unreviewed; the claimed reversal of losses and outperformance of standard multi-sample inference remain to be independently validated.
Preprint. Language models on reasoning benchmarks; no human subjects identified.. Intervention: Deformation-controlled, support-preserving Power Sampling target with weighted self-consistency aggregation.. Compared with: Standard Power Sampling with uniform trajectory exponentiation and standard multi-sample inference..
Power Sampling can drive probability mass toward correct trajectories while degrading downstream inference accuracy by up to 18.5 percentage points across models and reasoning benchmarks. Two mechanisms are identified: dose mismatch (fixed exponent induces different distributional change across problems) and coverage mismatch (global sharpening narrows support despite high pass@k diversity metric). A repaired sampler using deformation-controlled, support-preserving Power target with weighted self-consistency reverses losses and outperforms standard multi-sample inference.
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is an unreviewed preprint describing a methodological problem in language model sampling and proposing a fix, demonstrated empirically across benchmarks but without peer review or clinical/real-world validation.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Power Sampling sharpens a language model's distribution over complete generation trajectories, offering a verifier-free way to improve reasoning at inference time. It also has the potential to serve as a general-purpose front end for a broad range of downstream sampling methods. However, we uncover a striking paradox: Power Sampling can drive more probability mass toward correct trajectories while degrading the downstream inference it is intended to enhance. Using self-consistency as a representative case, we observe accuracy drops of up to 18.5 percentage points across models and reasoning benchmarks. We trace this paradox to two mismatches. Dose mismatch arises because a fixed exponent induces drastically different amounts of distributional change across problems. Coverage mismatch arises because global sharpening concentrates mass on a narrow set of dominant paths: high pass@k, often interpreted as evidence of preserved diversity, can therefore coexist with the loss of broad reasoning-path support required for downstream aggregation, search, and selection. Guided by this diagnosis, we replace uniform trajectory exponentiation with a deformation-controlled, support-preserving Power target that calibrates sharpening across problems while limiting the suppression of moderate-probability paths. In a same-budget instantiation with weighted self-consistency, the repaired sampler reverses the losses caused by global Power and outperforms standard multi-sample inference across reasoning benchmarks.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.