Life sciences · Preprint
arXiv · September 9, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint empirically demonstrates that temporal hold-outs (test data published after model release) do not eliminate domain-specific pretraining advantage in time-series foundation models. Pretrained models win on 5 of 7 dataset groups, but the advantage is driven by corpus familiarity (e.g., 28% lower MASE on Wikipedia pageviews, a domain described as bulk of TimesFM's pretraining corpus) rather than intrinsic time-series properties such as seasonal strength or spectral entropy. The work identifies a methodological gap in benchmarking practices but does not propose a validated solution.
Comparative benchmark study with temporal hold-out. Seven time-series dataset groups from five domains (electricity, finance, traffic, web analytics, weather/climate), each dataset rebuildable without API keys. All observations timestamped after the final model release.. Intervention: Temporal hold-out (test data from after model release date) to evaluate whether pretrained time-series foundation models retain domain familiarity advantage when memorisation of a specific time window is prevented.. Compared with: Classical statistical methods (e.g., Theta), per-dataset-trained neural networks, seasonal naive baseline, and competing pretrained foundation models (TimesFM, Chronos families).. Not stated..
Pretrained models outperformed classical baselines on 5 of 7 dataset groups; lost one to Theta baseline; indistinguishable from seasonal naive on daily exchange rates. Largest documented gain: 28% lower MASE than best classical method on weekly Wikipedia pageviews (the domain described as bulk of TimesFM's pretraining corpus). Within-family comparison: TimesFM family outranks Chronos family by −0.53 ranks on Wikipedia (1,500 series) versus −0.09 elsewhere (754 series), Mann-Whitney p < 1e-5.
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
This work is not a clinical study; it addresses machine learning practitioners and benchmark designers. The implications are methodological: temporal hold-outs alone are insufficient to prevent domain-familiarity bias in foundation model evaluation. Practitioners should assess whether their domain overlaps with the disclosed pretraining corpus before assuming a pretrained model will outperform domain-specific alternatives.
A methodologically sound empirical study of pretraining data leakage in time-series models, but limited to a single hold-out dataset, lacks peer review, and reports an important negative result about what does not predict performance without establishing causal mechanisms.
As stated by the source record.
Quoted from the source exactly as published.
This work is not a clinical study; it addresses machine learning practitioners and benchmark designers. The implications are methodological: temporal hold-outs alone are insufficient to prevent domain-familiarity bias in foundation model evaluation. Practitioners should assess whether their domain overlaps with the disclosed pretraining corpus before assuming a pretrained model will outperform domain-specific alternatives.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Time-series foundation models are evaluated almost exclusively on public archives that predate them, so a strong score cannot be separated from having seen the test set during pretraining. The obvious remedy is a hold-out that postdates the models. We build one: thirteen forecasters -- four classical, three trained per dataset, six pretrained -- on seven groups drawn from five domains, every observation published after the last model was released, and every dataset rebuildable without an API key. Under this protocol pretrained models win 5 of 7 groups, lose one to a Theta baseline, and on daily exchange rates are indistinguishable from a seasonal naive forecast, along with every other method tested. We then ask what separates the wins from the losses, and report a negative result: the two intrinsic properties one would reach for -- seasonal strength and spectral entropy, measured on the input window -- do not account for the pattern, and seasonal strength is if anything negatively associated with the advantage. What does track it is corpus familiarity. Our largest gain (28% lower MASE than the best classical method, on weekly Wikipedia pageviews) falls on Wikipedia pageviews, the domain TimesFM's authors describe as the bulk of its pretraining corpus, at the same granularities and differing only in time window. Within the pretrained family, where every model forecasts identical series so that series difficulty cancels, the TimesFM family outranks the Chronos family by -0.53 ranks on Wikipedia against -0.09 everywhere else (1,500 vs. 754 series, Mann-Whitney p < 1e-5). We conclude that a temporal hold-out removes memorisation of a window but not familiarity with a domain, that benchmarks therefore need domain hold-outs stated relative to disclosed corpora, and that the practitioner's question is less which model is better than whether their domain is one the model was raised on.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.