arXiv:2604.27405cs.CLcs.AI2026-04被引 3

用心理评估方法检测大模型版本间真实变化,发现多数更新效果被平均值掩盖。

Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation

论文配图:Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation
图 1 · 摘自论文原文
  • 引入临床心理学的可靠变化指数,评估模型版本间个体项变化
  • 超半数题目无可靠变化,但高低难度题表现相反:低分题提升,高分题退步
  • 单次评估易漏检42%真实变化,建议同时报告准确率与变化率

我们将临床心理学中的可靠变化指数(RCI)应用于大语言模型版本对比,针对MMLU-Pro的2000个题目进行分析(每题采样10次,温度T=0.7)。测试了两组同家族模型对:Llama 3到3.1(+1.6分)和Qwen 2.5到3(+2.8分)。在全量基准上,79%和72%的题目未显示可靠变化。但超过一半题目处于分数极值区(地板/天花板)。在可分析题目中,变化双向显著:Llama有34%改善、28%退步,Qwen则分别为47%和39%,中位绝对变化量达0.50和0.90。变化具有难度依赖性:低准确率题目改善,高准确率题目退步。领域分解显示家族特异性逆转:Llama在物理领域退步,而Qwen在法律领域退步。贪心单次评估遗漏42%可靠变化,误报25%无变化项目。整体准确率增长实为项级反向变动的净结果。建议报告变化率以补充整体准确率。

原文摘要 · Abstract (English)

We adapted the Reliable Change Index (RCI; Jacobson and Truax, 1991) from clinical psychology to item-level LLM version comparison on 2,000 MMLU-Pro items (K=10 samples at T=0.7). Two within-family pairs were tested: Llama 3 to 3.1 (+1.6 points) and Qwen 2.5 to 3 (+2.8 points). On the full benchmark, most items showed no reliable change (79% and 72%). However, over half the items were floor/ceiling. Among analysable items, change was bidirectional with large effect sizes: 34% improved and 28% deteriorated for Llama; 47% improved and 39% deteriorated for Qwen (median |delta p| = 0.50 and 0.90). Churn was asymmetric by difficulty: low-accuracy items improved, high-accuracy items deteriorated. Domain-level decomposition revealed family-specific reversals: Llama lost physics while Qwen lost law. Greedy single-shot evaluation missed 42% of reliably changed items and falsely flagged 25% of unchanged items. The aggregate accuracy gain is the net residual of opposing item-level movements. We recommend reporting churn rate alongside aggregate accuracy.

模型评估变化检测可靠性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。