arXiv:2608.17719cs.SEcs.AI2026-08

商用大模型升级时,平均分数掩盖了部分任务反而变差的真相。

What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations

  • 逐项分析900个测试题在3次版本迭代中的表现变化
  • 7.3%提升的版本中仍有8.3%任务明显退步,10.7%任务反而进步
  • 严格评分下退步3.9分,宽松评分下仅退0.04分,差异显著

依赖商用大语言模型API的软件系统在旧模型停用时需迁移至新版本。当前决策多依赖聚合基准分数,将异质的项目级行为压缩为单一数值。本文在GPT-5.4至GPT-5.6的三次版本升级中,对900个公开测试项(研究生级知识、奥数题、指令遵循)每项各测试50次,采用错误发现率控制与实际显著性阈值,分类每项为可靠提升、可靠退步、基本等效或不确定,并通过标签置换零模型校准结果。在全部九个迁移-基准组合中,可靠提升与可靠退步同时存在:聚合得分最高提升7.3个百分点的版本中,仍含高达8.3%的可靠退化项;聚合得分下降的版本中,最多有10.7%的可靠提升项。在指令遵循基准上,严格与宽松评分间的差距在最新迁移中扩大3.9个百分点——严格评分下退步3.9分,宽松评分下仅退0.04分。结论:仅凭聚合分数做迁移决策会忽略显著的双向项目级变化。完整响应级数据集与逐项评分输出已公开。

原文摘要 · Abstract (English)

Context: Software systems that depend on commercial large language model APIs must migrate to successor versions when vendors deprecate older models. Migration decisions typically rely on aggregate benchmark scores, which compress heterogeneous item-level behaviour into a single net figure. Objective: We measure what that compression conceals. Method: On three pairwise upgrades in the GPT-5.4 to GPT-5.6 Sol product sequence, we query 900 public benchmark items (graduate-level knowledge, olympiad mathematics, instruction following) 50 times per item per model, classify each item as reliably improved, reliably regressed, practically equivalent, or inconclusive under false-discovery-rate control and a practical-significance threshold, and calibrate the results against a label-permutation null. Results: Across all nine migration-benchmark cells, reliable improvements and reliable regressions coexist. Edges with aggregate gains of up to 7.3 percentage points contain up to 8.3% reliably regressed items; edges with aggregate losses contain up to 10.7% reliably improved items. On the instruction-following benchmark, the gap between strict and loose scoring widens by 3.9 percentage points on the latest migration: a 3.9-point regression under strict scoring shrinks to 0.04 points under loose scoring. Conclusion: Migration decisions based on aggregate scores alone miss substantial bidirectional item-level change. The complete response-level archive and per-item scoring outputs are released.

大模型评估迁移风险细粒度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。