Life sciences · Preprint
arXiv · September 9, 2026
Raises a question worth testing. It does not answer one.
This preprint introduces a 955-case benchmark for evaluating language models as transit kiosk policy engines, covering routing, fares, disruptions, and accessibility across six real metro systems. Performance is reported on deterministic (Tier 1) and semantic-quality (Tier 2) scoring; a 4B Qwen 3.5 student model with parameter-efficient fine-tuning achieves 91.3% on Tier 1, exceeding larger models and GPT variants. The benchmark is technical and reproducible but is not a clinical or operational trial; it does not measure real-world deployment success, user safety, or system reliability.
Benchmark evaluation with held-out test partition. Language model policy layers evaluated on transit kiosk scenarios covering six real metro systems (37–414 stations each), eleven task categories (routing, fare calculation, disruptions, accessibility, adversarial input, and others).. Intervention: Parameter-efficient fine-tuning (PEFT) of language models; maximum reasoning effort settings; different serving configurations.. Compared with: Base models (unpruned Qwen variants), GPT-5.6 Tiers 1 and 2, GPT-5.4 full, Muse Glimmer 30B, and a deterministic rule-based baseline.. Six real metro systems (specific locations not named in abstract)..
4B Qwen 3.5 student (PEFT) achieved 91.3% on Tier 1, exceeding GPT-5.6 tiers (90.6 and 90.0) and matching GPT-5.4 full at maximum reasoning effort (91.4) Footprint of 4B Qwen model is 2.6 GB Q4_K_M PEFT gain over base model ranges from +7.03 points at 2B to −0.91 at 27B across four Qwen sizes
Benchmark is simulation-based; does not measure real-world operational deployment, user satisfaction, safety outcomes, or actual system reliability.
The source did not state who this applies to in practice.
This is a technical benchmark paper evaluating language models on a simulated task; it reports comparative performance metrics but lacks clinical outcomes, real-world deployment data, or evidence of impact on actual transit systems or user safety.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, fare calculation, disruptions, accessibility, and adversarial input. In each case, the model must call structured tools and submit a machine-renderable terminal state containing an outcome, a per-ticket fare quote when applicable, and a kiosk action. Fourteen deterministic scoring components form Tier 1; eight semantic-quality components form Tier 2, six of which use a language-model judge. We report Tier 1 and the combined score of both tiers. A stratified 75/25 split reserves 717 cases for training-data generation and 238 for held-out evaluation. We evaluate twenty-six models from six vendors, of which twenty-three are ranked. On the held-out partition, a 4B Qwen 3.5 student trained through parameter-efficient fine-tuning (PEFT) exceeds both GPT-5.6 tiers on Tier 1 (91.3 against 90.6 and 90.0) and matches GPT-5.4 full at maximum reasoning effort (91.4), with a 2.6 GB Q4_K_M footprint. Larger 9B and 27B students provide no further Tier 1 improvement over the 4B student at this training scale. Across the four Qwen sizes, the PEFT gain over the corresponding base model decreases from +7.03 points at 2B (three training seeds) to -0.91 at 27B; every seed shows the same direction at every size. A deterministic rule-based baseline reaches 84.6 on Tier 1, with the remaining language-model advantage concentrated in policy adaptation, compound scenarios, accessibility, and temporal reasoning. Muse Glimmer 30B leads the composite ranking, and serving configuration alone moves the Qwen 3.5-to-3.8 comparison by 2.7 Tier 1 points. The benchmark, harness, reproduction guide, and fine-tuned students are released at https://github.com/continker/metrollm-bench.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.