arXiv:2608.11027cs.LGcs.CL2026-08

用统一测试集分析大模型行为演化,发现模型越来越相似且推理型模型更稳定。

Mapping and Measuring the Behavioral Evolution of Large Language Models

论文配图:Mapping and Measuring the Behavioral Evolution of Large Language Models
图 1 · 摘自论文原文
  • 用32个模型对1万条提示的回应构建句子级差异度量
  • 模型家族聚类清晰,跨家族距离随时间缩小,近期推理模型响应更紧凑
  • 方法无需标签,可压缩至原大小1/73仍保持关键结构

基准排行榜反映模型性能,但无法揭示其行为与其他模型的关系或代际变化。本文利用同一组10,000条提示,分析6大模型家族共32个模型的输出行为。通过嵌入每个响应,构建三种互补的句级差异指标:每提示的对齐均值距离(伪度量)、按提示分歧的PCA压缩摘要,以及模型内部响应几何结构间的无对齐Gromov–Wasserstein差异。基于这些指标,绘制行为图谱,分析静态组织与时间演化,包括家族漂移、层次聚类、跨家族趋同和响应云扩散。结果显示,模型家族形成连贯聚类,gpt-2为全局异常点;跨家族距离随时间减小;多个近期推理导向模型响应云更紧凑。基于令牌级别的最大均值差异(MMD)的交叉验证与句级均值距离高度一致(斯皮尔曼ρ=0.98),复现相同定性结论。方法采用测度论视角,明确对齐与不变性假设。同时提出一个架构无关的充分条件,将行为相似性与推理-提示覆盖范围、小过剩群体对数损失及相似有效目标分布关联起来,提供可能的训练侧解释而非仅经验现象。整个流程无需标签,即使将响应重新编码为三个额外编码器(最小至原大小1/73),仍保留排名几何结构、异常点和时间趋势符号。

原文摘要 · Abstract (English)

Benchmark leaderboards summarize how well a language model performs, but not how its behavior relates to that of other models or changes across generations. We characterize the output behavior of 32 models from six families using their responses to a shared bank of 10{,}000 prompts. After embedding each response, we construct three complementary sentence-level dissimilarities: an aligned mean per-prompt distance, which is a pseudometric on observed model responses; a PCA-compressed summary of prompt-wise disagreement; and an alignment-free Gromov--Wasserstein discrepancy between models' internal response geometries. We use these constructions to study static organization and temporal change on a release-date axis through behavioral maps, family-wise drift, hierarchical clustering, cross-family convergence, and response-cloud dispersion. Across the three constructions, model families form coherent clusters, with \texttt{gpt-2} as a global outlier; cross-family distances decrease over time; and several recent reasoning-oriented models have comparatively compact response clouds. A token-level cross-check based on per-prompt Maximum Mean Discrepancy closely agrees with the sentence-level mean distance (Spearman $ρ=0.98$) and recovers the same qualitative findings. We organize these comparisons through a measure-theoretic lens making their alignment and invariance assumptions explicit. We also establish an architecture-agnostic sufficient condition linking behavioral similarity to inference-prompt coverage, small excess population log-loss, and similar effective target distributions---a possible training-side account rather than an empirical explanation of the observed trends. Our pipeline is label-free, and re-encoding every response with three further encoders---down to one $73\times$ smaller---preserves the rank geometry, the outliers, and the sign of the time trend.

大模型行为模型演化相似性度量无监督分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。