Life sciences · Preprint
arXiv · September 8, 2026
Early or partial results. Treat as a signal, not a conclusion.
This is a controlled human study comparing six reasoning representation formats on human ability to evaluate large language model outputs. The study reports a mismatch between participant preferences and task performance: planning and decomposition formats were preferred, but simpler chain-of-thought traces performed better on verification, trust, and interpretability tasks. The work raises questions about the design of human-facing AI explanations but lacks sufficient detail on sample size, population characteristics, and statistical analysis to support strong claims.
Controlled human study with randomization. Study participants evaluating LLM outputs; specific eligibility criteria, recruitment method, and participant count not stated.. Intervention: Six reasoning representation formats for LLM outputs, including planning-, decomposition-, and chain-of-thought-based representations.. Compared with: Comparison across six reasoning formats; no external control or baseline stated..
Participants prefer planning- and decomposition-based representations despite their lower support for human evaluation tasks. Simpler chain-of-thought traces better support verification, trust, and interpretability compared to preferred formats. Preferred representations introduce calibration risks, with more false alarms on correct traces and high trust despite low willingness to verify.
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
A single-centre controlled human study of reasoning representations with qualitative and judgement-based outcomes, not a clinical or hard endpoint trial; findings are exploratory and require replication.
As stated by the source record.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
What is missing. This record has no reported figures. That is a gap in the analysis, not a judgement about the study.
Reasoning representations are increasingly used as explanations for large language model outputs. Yet they are typically evaluated with model-centric criteria, such as answer accuracy and faithfulness, leaving it unclear whether they help people evaluate model responses. In this work, we study reasoning representations as human-facing interfaces rather than proxies for model reasoning ability. We conduct a controlled human study of six reasoning formats across tasks of varying complexity, supported by a web-based framework that randomizes task domains, problem instances, and representation order. The study collects fine-grained judgments of structural understanding, error detection and localization, and trust calibration. Our study shows a mismatch between perceived preference and support for human evaluation. Participants prefer planning- and decomposition-based representations, but simpler chain-of-thought traces better support verification, trust, and interpretability. Preferred representations also introduce calibration risks, with more false alarms on correct traces and high trust despite low willingness to verify.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.