Life sciences · Preprint
arXiv · September 9, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint reports a computational study testing whether language models possess accurate self-knowledge about their own behavior. The authors find that direct self-report correlates very weakly with measured behavior (r = +0.04), and that generic statements about AI agents in general predict a model's own behavior almost as well as the model's self-reports. The study indicates that self-descriptions in language models reflect a generic theory of AI assistants and a favorable bias rather than privileged introspective knowledge.
Comparative computational study with control conditions. Language models, including frontier-scale variants; evaluated on nine behavioral domains including resistance to pushback, tool misuse, and deception under pressure.. Intervention: Self-report prediction task; item-informed self-prediction; finetuning on behavioral record.. Compared with: Generic AI agent predictions, other models' self-predictions, measured actual behavior..
Direct self-report correlation with actual behavior is weak (r = +0.04), rising only to +0.24 when models see exact items. Generic questions about capable AI agents predict target model behavior (+0.28) at least as well as the model's own self-predictions. Other models' answers about themselves predict target model behavior as well as the target model's own answers.
First-person framing consistently shifts reports in a flattering direction, understating harmful behavior relative to generic agent questions.
The source did not state who this applies to in practice.
Exploratory computational study testing self-knowledge in language models using novel methodology, with clear quantitative results but limited scope, no peer review, and no direct clinical or practice guidance.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Language models can fluently describe how they would behave: whether they would cave to pushback, misuse a tool, or lie under pressure. Is that description actually about the model speaking? We turn self-knowledge into a prediction test. Across nine behavioral evaluations, we measure how a model behaves under different conditions, ask it to predict those rates, and compare its predictions with controls that remove the self from the question. We find that: (i) Direct self-report is weak (r = +0.04), and even showing the model the exact items only raises prediction to +0.24. Crucially, the same item-informed question about "capable AI agents in general" does just as well (+0.28), while other models' answers about themselves predict the target model at least as well as its own. (ii) Frontier scale does not detectably change this pattern: any gains in prediction are not self-specific, and are consistent with a better theory of how AI assistants behave rather than better self-knowledge. (iii) First-person framing does have one robust effect: it shifts reports in the flattering direction, understating harmful behavior relative to the same question about a generic agent. (iv) Finetuning on a model's own behavioral record can teach narrow self-predictions, but it also changes the behavior being predicted and the gains do not transfer broadly. The practical implication is simple: asking a model what it would do mostly reveals a theory of AI assistants in general, plus a favorable bias, rather than privileged knowledge of that model.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.