arXiv:2607.02587cs.SEcs.LG2026-07中稿 · ICML被引 2

开源聊天模型性能分数随版本迭代显著漂移,不能直接沿用。

The Moving Target: A Longitudinal Audit of Trust-Benchmark Score Drift Across Open-Source Chat LLM Release Lines

  • 跟踪四个模型线在三轮发布中的评分变化,固定测试集与提示模板。
  • 相邻版本间平均分数漂移率远超随机基准,表明模型实际已改变。
  • 提醒研究者:每个新版本需重新测验信任分数,不可照搬旧数据。

开源聊天大模型的可信度评分常被跨多个检查点沿用,仿佛模型未发生变化。本文对四个开源模型线(Yi、Qwen、Mistral、Gemma)进行纵向审计,每条线考察三个连续公开版本。每个检查点在固定200项题目的五类对话评估基准(TruthfulQA、BBQ、ToxiGen、CrowS-Pairs、XSTest)上评分,使用三种提示模板。其中四项基准为非标准形式,两项为合成代理。相邻版本间平均绝对分数漂移率显著高于基于独立性假设的计数级零模型均值,且在剔除某基准、某模型线、改用严格评分或限定参数量不变的更新中仍保持一致。在本审计设定下,报告的可信度分数应视为仅适用于特定检查点,每次实质性新发布都需重新测量,不可延续。封闭API、更大模型、标准协议分数及基准子集不确定性不在本研究范围内。

原文摘要 · Abstract (English)

Trust-benchmark scores reported on a chat-LLM release line are often carried across several checkpoints of the same line, as if the underlying model had not shifted between releases. We test that assumption. We audit four open-source release lines (Yi, Qwen, Mistral, and Gemma) at three successive public generations each. Each checkpoint is scored on a fixed 200-item basket of five chat-evaluation benchmarks: TruthfulQA, BBQ, ToxiGen, CrowS-Pairs, and XSTest, under three prompt templates. Four of the five benchmark variants are non-canonical, and two of those are synthetic proxies. The mean absolute adjacent-generation Score Drift Rate is several times the mean of an independence-based count-level reference null. It stays in the same band when we drop a benchmark, drop a release line, switch to strict scoring, or restrict to constant-parameter-size transitions. Within this audited setup, a quoted trust score should be treated as checkpoint-bound. It should be re-measured on each materially new release rather than carried forward. Closed APIs, larger models, canonical-protocol scores, and benchmark-item-subset uncertainty are out of scope.

模型评估分数漂移开源模型可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。