Life sciences · Preprint
arXiv · August 17, 2026
Posted before peer review. The findings may change or fail to hold.
This is an unpublished preprint describing two proprietary large language models (Mint-Cu 9B and Mint-Ag 27B) designed for financial tasks. The authors report superior performance on three internal or proprietary benchmarks compared to named competitors, but provide no peer review, independent validation, or evidence of real-world utility in clinical or financial decision-making.
Preprint. Intervention: Two proprietary large language models: Mint-Cu (9B parameters) and Mint-Ag (27B parameters), trained using SFT, critical-step OPD, RLVR, model merging, and multi-teacher on-policy distillation.. Compared with: GPT-5.6-Sol, Claude-Opus-4.8, Agents-A1-35B, and Nex-N2-mini on stated benchmarks..
Mint-Ag achieves 98.33% on RFC-Bench, reported as 3.66 points above GPT-5.6-Sol and 3.00 points above Claude-Opus-4.8 Mint-Cu reaches 69.86% on FinSearchComp T2, outperforming Agents-A1-35B by 22.83 points and Nex-N2-mini by 12.78 points Mint-Ag achieves 76.00% on FinanceAgentBench v1.1 and 60.49% on FinanceAgentBench v2
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is an unrefereed arXiv preprint describing model development and benchmark performance; it reports no clinical outcomes, human trials, or peer-reviewed validation, and addresses computational rather than medical evidence.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Financial agents must do more than recall domain knowledge: they must be both reliable, executing precise operations over grounded evidence, and executive, sustaining long-horizon research whose conclusions remain auditable. We present Mint-Agent, a family of finance-native agentic models designed around these two scales of financial intelligence. Mint-Agent is built upon three pillars: data, harness, and algorithm. Our data engine constructs clean, specialized tasks for atomic financial capabilities and long-horizon agentic execution from real-world financial sources. MintHarness enables stable interaction with open-ended environments and maintains auditable evidence trails across extended research trajectories. Our training recipe combines SFT, critical-step OPD, and RLVR to develop separate financial reasoning and agentic execution experts, which are then unified through model merging and multi-teacher on-policy distillation into compact, general-purpose financial agents. This pipeline yields two flagship models, Mint-Cu (9B) and Mint-Ag (27B). Across professional financial benchmarks, our models demonstrate two defining strengths: (1) Reliability: Mint-Ag achieves 98.33% on RFC-Bench, surpassing GPT-5.6-Sol and Claude-Opus-4.8 by 3.66 and 3.00 points; and (2) Executability: Mint-Cu reaches 69.86% on FinSearchComp T2, outperforming Agents-A1-35B and Nex-N2-mini by 22.83 and 12.78 points, while Mint-Ag achieves 76.00% and 60.49% on FinanceAgentBench v1.1 and v2, respectively. These results establish a path toward trustworthy financial intelligence in which domain expertise, long-horizon execution, and auditable evidence are jointly engineered as a unified foundation for frontier agentic models.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.