用新方法诊断大模型能力协同关系,帮团队决定该调模型还是等下一轮。
The Growing Pains of Frontier Models: When Leaderboards Stop Separating and What to Measure Next

- 拆解双基准分数,分离出模型间能力协同趋势与每轮独立表现差异
- 34个模型显示能力普遍协同(相关系数0.72),但不同实验室差异达5倍
- 提供可操作的诊断三步法,指导何时重训、何时等待或调整推理算力
排行榜在独立维度上排名前沿模型,却无法揭示能力之间是相互促进还是此消彼长——而在前沿阶段,这种交互关系才是更关键的信号。我们通过将配对的 SWE-bench 与 GPQA Diamond 分数分解为群体耦合趋势和每轮残差(h-场),诊断出多个公开基准下的能力侧重。覆盖10家实验室共34个模型(2024–2026年发布),发现能力普遍协同(r = +0.72, p < 10⁻⁶),但协同程度系统性差异显著:各实验室耦合斜率跨度达5倍(谷歌1.15,DeepSeek 0.23),且部分实验室发生转向——DeepSeek从以推理为主转为以编码为主(Δh = 15.9 pp);Anthropic则在编码突进与恢复间周期震荡。群体回归作为等值线相位边界,同一√((a/b)·B₁)分类器不仅识别基础规模耦合转变[ Amin, 2026 ],还能提前检测下一阶段的混合相行为(两个模型已低于GPQA–IFEval等值线)。h-场不仅是诊断工具,更指引改进方向:预训练确立耦合水平为0.871,而强化学习微调(RLHF)带来+0.081提升[ Amin, 2026 ]——预训练级变化不可逆(如DeepSeek四轮转向持续存在),后训练变化可逆(Anthropic三次编码激增均在一版内恢复),仅增加推理算力即可使h提升+7.8 pp而无需重训。掌握主导成分可决定是否重训或等待。本文提供三步诊断流程(定位、分类、预测)、各实验室测量优先级表,以及七条可验证的预测(带时间戳标准)。五次截断后的发布结果均落在95%预测区间内。代码、数据与交互式仪表盘见:https://zehenlabs.com/cape/
原文摘要 · Abstract (English)
Leaderboards rank frontier models on independent axes but do not reveal whether capabilities reinforce or trade off across releases -- and at the frontier, this interaction is the more informative signal. We decompose paired SWE-bench and GPQA Diamond scores into a population coupling trend and per-release residual ($h$-field) that diagnoses capability emphasis from two public benchmark scores. Across 34 models from 10 labs (2024--2026), capabilities cooperate ($r = +0.72$, $p < 10^{-6}$), but cooperation varies systematically: per-lab coupling slopes span $5\times$ (Google $1.15$ vs. DeepSeek $0.23$), and labs pivot -- DeepSeek reversed from reasoning-rich to coding-first ($Δh = 15.9$~pp); Anthropic oscillates between coding excursions and recovery. The population regression serves as an isocline phase boundary: the same $\sqrt{(a/b)\cdot B_1}$ classifier that identifies the base-scale coupling transition [Amin, 2026] classifies frontier models and already detects mixed-phase behavior at the next transition (two models below the GPQA--IFEval isocline). The $h$-field is not just diagnostic -- it tells you what to change. Pretraining establishes coupling at $0.871$ while RLHF adds $0.081$ [Amin, 2026]: pretraining-level shifts are permanent (DeepSeek's four-release reversal persists), post-training shifts are reversible (Anthropic's three coding excursions each recover within one release), and inference compute alone shifts $h$ by $+7.8$~pp without retraining. Knowing which component dominates determines whether to retrain or wait. We provide a three-step diagnostic (locate, classify, predict), a per-lab measurement-priority table, and seven falsifiable predictions with timestamped criteria. Five post-cutoff releases fall within the 95\% prediction interval. Code, data, and an interactive dashboard: https://zehenlabs.com/cape/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。