Life sciences · Preprint
arXiv · August 13, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint describes development of Faraday, a 27-billion-parameter AI agent trained on a custom paper replication task space using an auto-generated rubric-based evaluator. The work reports that the agent outperforms Claude Opus 4.8 and GPT-5.5 on held-out replication tasks and adopts a more principled approach in qualitative analysis, but lacks peer review, independent validation, reporting of quantitative performance metrics, and demonstration of utility for actual scientific practice.
Uncontrolled AI system development and benchmarking study. Intervention: Faraday, a 27B-parameter AI Scientist agent trained on paper replication tasks using a custom rubric-based judge. Compared with: Claude Opus 4.8 and GPT-5.5 on held-out replication tasks.
Faraday, a 27B-parameter AI Scientist agent, surpasses performance of Claude Opus 4.8 and GPT-5.5 on held-out replication tasks Auto-generated rubric-based judge has low noise and agrees with human assessment of replication quality Qualitative analysis reveals Faraday adopts a more scientifically-principled approach than comparison models
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is an early-stage, uncontrolled development report of an AI system for research replication, without peer review, clinical outcomes, or comparison to established human baselines in a prospective trial.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
The replicability of papers is a cornerstone of scientific knowledge, ensuring the reliability of existing results and providing a base for further experiments. The act of replication typically illuminates details that were previously underspecified, and thus requires similar hypothesis-driven exploration to open-ended research. In this work, we develop Replica, a scalable task space for paper replication. To provide reward signal, we introduce an auto-generated rubric-based judge that has low noise and agrees with human assessment of replication quality. We post-train Faraday, a 27B-parameter "AI Scientist" agent that leverages coding agents as tools, surpassing the performance of Claude Opus 4.8 and GPT-5.5 on held-out replication tasks. Qualitative analysis of individual rollouts reveals that Faraday adopts a more scientifically-principled approach. We believe that our results provide a stepping stone towards AI agents capable of long-horizon scientific innovation without requiring complex harnesses.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.