Life sciences · Preprint
arXiv · August 11, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint presents an unsupervised, measure-theoretic framework for characterizing and comparing the behavioral similarity of large language models across families and over time, using multiple geometric distance metrics on model responses. Key findings include coherent clustering within model families, decreasing cross-family distances over time, and recovery of consistent patterns across encoding schemes. The work is methodological and exploratory; it does not establish clinical utility, causal mechanisms, or predictive validity.
Observational comparative analysis of model outputs using unsupervised embedding and geometric dissimilarity metrics. 32 large language models from six families. No specification of model sizes, training data, or other characteristics beyond family membership and release date.. Compared with: Pairwise comparison of behavioral outputs across models and families using three geometric dissimilarity constructions.. n = 32.
Model families form coherent clusters with gpt-2 identified as a global outlier across three dissimilarity constructions Cross-family distances decrease over time, indicating behavioral convergence across recent model generations Token-level Maximum Mean Discrepancy cross-check shows high agreement with sentence-level mean distance (Spearman ρ=0.98) and recovers the same qualitative findings
No assessment of downstream task performance, safety, or alignment; behavioral similarity does not imply equivalent utility or safety.
The source did not state who this applies to in practice.
This is an unsupervised exploratory analysis characterizing behavioral similarities across language models using multiple distance metrics, without a clinical endpoint or validated predictive outcome; it establishes methodology and describes phenomena but does not establish causal mechanisms or clinical utility.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Benchmark leaderboards summarize how well a language model performs, but not how its behavior relates to that of other models or changes across generations. We characterize the output behavior of 32 models from six families using their responses to a shared bank of 10{,}000 prompts. After embedding each response, we construct three complementary sentence-level dissimilarities: an aligned mean per-prompt distance, which is a pseudometric on observed model responses; a PCA-compressed summary of prompt-wise disagreement; and an alignment-free Gromov--Wasserstein discrepancy between models' internal response geometries. We use these constructions to study static organization and temporal change on a release-date axis through behavioral maps, family-wise drift, hierarchical clustering, cross-family convergence, and response-cloud dispersion. Across the three constructions, model families form coherent clusters, with \texttt{gpt-2} as a global outlier; cross-family distances decrease over time; and several recent reasoning-oriented models have comparatively compact response clouds. A token-level cross-check based on per-prompt Maximum Mean Discrepancy closely agrees with the sentence-level mean distance (Spearman $ρ=0.98$) and recovers the same qualitative findings. We organize these comparisons through a measure-theoretic lens making their alignment and invariance assumptions explicit. We also establish an architecture-agnostic sufficient condition linking behavioral similarity to inference-prompt coverage, small excess population log-loss, and similar effective target distributions---a possible training-side account rather than an empirical explanation of the observed trends. Our pipeline is label-free, and re-encoding every response with three further encoders---down to one $73\times$ smaller---preserves the rank geometry, the outliers, and the sign of the time trend.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.