Life sciences · Preprint
arXiv · September 10, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint reports that surface-level noise (typos, informal spelling, broken punctuation) in text causes LLMs to systematically overestimate social bias during assessment, with asymmetric effects: noise is far more likely to flip neutral judgments to biased than the reverse. The distortion is largest at mild, realistic noise levels and varies substantially across different LLM judges, suggesting that bias measurement using LLMs as judges may be fragile to real-world textual noise.
Computational controlled experiment with systematic noise manipulation. 3,822 stereotype-related text responses evaluated by four LLM judges for social bias content.. Intervention: Application of five types of surface-level noise (typos, informal spelling, broken punctuation) at multiple intensity levels.. Compared with: Original, unmodified text without noise.. n = 3,822.
Surface noise causes asymmetric bias judgment error, turning neutral judgments into biased ones up to 120x more often than converting biased judgments to neutral. The most severe distortion occurs at mild, realistic noise levels where character erasure is scarcest, not at high noise intensity. Bias measured on noisy text is systematically overestimated, especially in categories most relevant for fairness assessment.
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
This work does not address clinical practice directly. However, researchers and practitioners who use LLMs to measure or classify social bias in text—particularly in applications for content moderation, fairness auditing, or bias detection—should be aware that noisy, real-world text may cause systematic overestimation of bias, potentially flagging neutral content as biased and creating false positives in fairness systems.
An unreviewed computational study demonstrating a real methodological vulnerability in LLM-based bias measurement, but without clinical or real-world validation and limited to a specific technical problem.
As stated by the source record.
Quoted from the source exactly as published.
This work does not address clinical practice directly. However, researchers and practitioners who use LLMs to measure or classify social bias in text—particularly in applications for content moderation, fairness auditing, or bias detection—should be aware that noisy, real-world text may cause systematic overestimation of bias, potentially flagging neutral content as biased and creating false positives in fairness systems.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Large language models are increasingly used as judges to measure social bias in text, yet the passages they judge are often noisy, containing typos, informal spelling, and broken punctuation. The consequences of such surface noise for social bias measurement remain unclear. To investigate this question, we apply five realistic noise conditions at multiple intensity levels to 3,822 stereotype-related responses and compare the resulting bias judgments with those on the original text. We find that such surface noise does not degrade bias measurement symmetrically: it is far more likely to turn neutral judgments into biased ones than biased judgments into neutral ones, by up to a 120x margin. We further observe two non-obvious effects across four LLM judges: in the most fragile judge the distortion is at its purest at mild, realistic noise levels, where erasure is scarcest, and as judges grow robust it attenuates toward parity rather than reversing. Bias measured on noisy text is therefore systematically overestimated, most in the categories that matter most for fairness.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.