Life sciences · Preprint
arXiv · August 12, 2026
Raises a question worth testing. It does not answer one.
This computational study uses pass@K and avg@K metrics to re-examine claims about on-policy distillation (OPD) in large language models. The authors report that OPD-trained models show better average-case performance but lose advantage in best-of-many settings, and argue this reflects improved sampling efficiency rather than genuine capability expansion—a reframing they term 'illusory distillation'.
Computational model evaluation study. Large language models before and after on-policy distillation training. Intervention: On-policy distillation (OPD) training on student models. Compared with: Pre-OPD base models, evaluated at varying sampling budgets.
OPD-trained models maintain superior avg@K performance across sampling budgets Advantage in pass@K gradually shifts to pre-OPD base models as K increases Pass@K dynamics show progressive shift toward stronger small-K performance at the expense of large-K capability boundary
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is a computational analysis using existing models to examine mechanisms underlying on-policy distillation; it raises questions about the interpretation of OPD rather than providing empirical evidence of clinical or practical utility.
As stated by the source record.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
What is missing. This record has no reported figures. That is a gap in the analysis, not a judgement about the study.
On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. It is commonly believed to enable the student model to distill knowledge from a stronger teacher model, thereby expanding capabilities beyond the pre-OPD base model. In this study, we examine this view through the lens of test-time scaling by varying the sampling budget K and evaluating performance with pass@K and avg@K. Specifically, across several OPD variants, we observe that OPD-trained models maintain superior avg@K performance across sampling budgets, while the advantage in pass@K gradually shifts to the pre-OPD base models as K increases. These results suggest that OPD primarily improves sampling efficiency rather than consistently expanding the student's reasoning capability boundary. The pass@K dynamics throughout OPD training further reveal a progressive shift toward stronger small-K performance at the expense of the large-K capability boundary. Furthermore, a problem-level solvability analysis using pass@1024 as the criterion reveals an asymmetry: OPD causes more previously solvable problems to become unsolvable than previously unsolvable problems to become solvable. Together, these findings suggest that, from the perspective of capability expansion, OPD behaves more like an "illusory distillation": its apparent gains arise primarily from improved sampling efficiency rather than from acquiring genuinely new reasoning capabilities from the teacher.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.