Life sciences · Preprint
arXiv · September 4, 2026
Raises a question worth testing. It does not answer one.
This preprint audits the Duckworth-Lewis-Stern method, the international standard for rain-revised cricket targets, on 8,150 historical matches and reports two systematic prediction biases: a 137-run range of error across match states and a gender-differential bias (+6.13 runs higher over-prediction for women's ODIs). The authors propose DLS-Cal, a calibration layer that reduces absolute bias by 31% on ODIs and 19% on T20Is, and a gender-aware variant that narrows women's residual bias from +6.19 to +0.65 runs. However, this is a computational and statistical study with no prospective validation, real-world fairness outcome data, or evidence of operational uptake.
Retrospective observational audit of historical international cricket matches with temporal train-test split; benchmarking study against multiple machine-learning comparators; model development without prospective validation. International limited-overs cricket matches from Cricsheet: men's and women's ODIs and T20Is. Specific eligibility criteria, inclusion/exclusion rules, and match-state distributions not detailed in abstract.. Intervention: DLS-Cal, a lightweight interpretable calibration layer (27K parameters) outputting a state-conditioned correction added to standard Duckworth-Lewis-Stern predictions; also a gender-aware variant of DLS-Cal.. Compared with: Duckworth-Lewis-Stern method (current international standard since 1999); also benchmarked against five modern machine-learning methods: Bi-LSTM, XGBoost, enriched XGBoost, deep context-aware model, and stacking ensemble.. n = 8,150. International matches (venue, dates, and geographic coverage not specified in abstract); data source is Cricsheet..
DLS prediction error spans 137-run range across match-state buckets (overs-remaining and wickets-lost combinations) Gender-differential bias on ODIs: mean over-prediction +1.51 runs for men versus +7.63 runs for women, gap of +6.13 runs (F = 195.16, p < 10^-43) DLS-Cal reduces absolute bias by 31% on ODI and 19% on T20I formats
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
This study is not clinical. For sports analytics professionals and cricket administrators, it provides quantified evidence of systematic prediction biases in the 25-year-old standard, particularly a previously undocumented gender-differential bias in women's ODIs. However, adoption of the proposed correction would require independent validation, stakeholder consensus, and evidence that correcting prediction error improves actual fairness in match outcomes—none of which this preprint addresses.
This is a computational audit and model development study on historical cricket data that identifies statistical patterns and proposes a calibration method, but lacks clinical or sports-outcome validation and does not constitute evidence that the proposed correction should be adopted operationally.
As stated by the source record.
Quoted from the source exactly as published.
This study is not clinical. For sports analytics professionals and cricket administrators, it provides quantified evidence of systematic prediction biases in the 25-year-old standard, particularly a previously undocumented gender-differential bias in women's ODIs. However, adoption of the proposed correction would require independent validation, stakeholder consensus, and evidence that correcting prediction error improves actual fairness in match outcomes—none of which this preprint addresses.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
The Duckworth-Lewis-Stern (DLS) method has been the international standard for revising target scores in rain-interrupted limited-overs cricket since 1999. Despite over two decades of operational use, no large-scale empirical audit of its prediction bias has been published. We conduct such an audit on 8,150 international matches (3,095 ODIs, 5,055 T20Is) from Cricsheet, generating 233,550 synthetic interruption scenarios with temporal splits. We document two structured biases. First, DLS prediction error spans a 137-run range across (overs-remaining, wickets-lost) match-state buckets. Second, DLS exhibits a gender-differential bias on ODIs that has not previously been quantified: on the training split, mean over-prediction is +1.51 runs for men but +7.63 runs for women, a gap of +6.13 runs (F = 195.16, p < 10^-43). We benchmark DLS against five modern alternatives: Bi-LSTM, XGBoost, an enriched XGBoost variant, a deep context-aware model, and a stacking ensemble, and propose DLS-Cal, a lightweight interpretable calibration layer (27K parameters) outputting a state-conditioned correction added to DLS. DLS-Cal reduces absolute bias by 31% on ODI and 19% on T20I, and a gender-aware variant reduces women's ODI residual bias from +6.19 to +0.65 runs while leaving men's calibration unchanged. We release code, models, and data.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.