Life sciences · Preprint
arXiv · September 10, 2026
Posted before peer review. The findings may change or fail to hold.
This preprint describes T1, a 122B-parameter Mixture-of-Experts model trained with reinforcement learning on long-horizon terminal tasks. The authors report improvements on Terminal-Bench 2.1 (43.8% to 64.0%) and performance on Long-Horizon Terminal Bench (27.9%), with comparison figures against GPT and GLM, but the work has not undergone peer review and lacks sufficient methodological detail and independent validation.
Preprint. Long-horizon terminal tasks including coding and scientific discovery; evaluation on Terminal-Bench 2.1 and Long-Horizon Terminal Bench. Intervention: T1 Mixture-of-Experts model (122B parameters) trained with reinforcement learning using dense process reward and terminal task verifier; TITO construction and R3 replay routing techniques applied. Compared with: Base model (initial 43.8% on Terminal-Bench 2.1); reported comparisons to GPT-5.4 and GLM-5.1 on Long-Horizon Terminal Bench. Cloud sandbox (specific location not stated).
T1 raised initial base model performance from 43.8% to 64.0% on Terminal-Bench 2.1 T1 achieved 27.9% on Long-Horizon Terminal Bench, reported to surpass GPT-5.4 and GLM-5.1 TITO and R3 techniques reduced training-to-inference log-probability difference from 0.021 to 0.013
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is an unrefereed arXiv preprint describing a machine learning model development and benchmark evaluation with no peer review reported; it reports empirical results but lacks independent validation.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Agent usage is shifting toward long-horizon tasks such as coding and scientific discovery, among which terminal tasks are especially important. We introduce T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task's own verifier. We provide a comprehensive recipe: First, an aggressively warm-started to stabilize actor-critic training, with a dense process reward scoring trajectories by the absolute number of passing verifiers. Second, stable optimization through TITO construction, training on the exact sampled token identifiers with drift repair at turn boundaries, and rollout routing replay, recording the sampler's per-token expert choices at every MoE layer and replaying them during training. Third, fully out-of-distribution training corpus: isolated seeds and synthesized tasks disjoint from Terminal-Bench 2.1 ensures gains reflect genuine capability transfer over benchmark overfitting. Together, TITO and R3 cut the training-to-inference log-probability difference from 0.021 to 0.013, with exactly aligned zero token drift in the loss region. On Terminal-Bench 2.1, our post-train pipeline raises initial base model from 43.8% to T1 with 64.0% resolved. On Long-Horizon Terminal Bench, T1 reaches 27.9% and surpasses GPT-5.4 and GLM-5.1.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.