Life sciences · Preprint
arXiv · September 8, 2026
Posted before peer review. The findings may change or fail to hold.
This preprint proposes RouteOPD, a reformulation of on-policy distillation as probability transport with explicit routing of student probability mass to teacher-preferred destinations. Experiments show consistent outperformance over reverse-KL baselines across four mathematical-reasoning benchmarks. However, the work has not undergone peer review, lacks statistical rigor reporting, and has no direct bearing on clinical practice.
Preprint. Intervention: RouteOPD (Routed On-Policy Distillation): a probability transport formulation decomposing teacher-student disagreement into student-excess sources and teacher-deficit destinations coupled into explicit transport pairs. Compared with: Sampled reverse-KL on-policy distillation.
RouteOPD consistently outperforms sampled reverse-KL OPD across four teacher-student settings Improvements accompanied by higher routing fidelity and lower background leakage Method tested on four mathematical-reasoning benchmarks
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is an unrefereed arXiv preprint proposing a novel machine learning method with experimental validation across benchmarks, but lacks peer review and is not a clinical or medical evidence study.
As stated by the source record.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
What is missing. This record has no reported figures. That is a gap in the analysis, not a judgement about the study.
On-policy distillation (OPD) transfers teacher knowledge on student-generated trajectories, but efficient sampled objectives reduce the teacher distribution to scalar credit on individual tokens. Such credit indicates whether a token should gain or lose probability, yet leaves the corresponding redistribution unspecified. We recast OPD as teacher-guided probability transport and propose RouteOPD (Routed On-Policy Distillation), which decomposes local teacher--student disagreement into student-excess sources and teacher-deficit destinations and couples them into explicit transport pairs. RouteOPD optimizes pairwise log-odds toward jointly realizable targets obtained from a bounded teacher potential, while adapting the transport budget to the concentration of teacher demand. This formulation directs updates toward teacher-preferred destinations and controls their magnitude within a single transport operator. Experiments across four teacher--student settings and four mathematical-reasoning benchmarks demonstrate that RouteOPD consistently outperforms sampled reverse-KL OPD, with improvements accompanied by higher routing fidelity and lower background leakage. These results demonstrate the effectiveness of explicitly modeling probability transport in on-policy distillation.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.