API评测结果不能可靠反映聊天机器人实际表现。
API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

- 对比API与界面接口的模型表现,发现系统性差异。
- 平均准确率差3.4个百分点,一致性差2.1个百分点。
- 适合关注模型落地真实效果的研究者和决策者。
基准测试分数是模型发布中的核心指标:影响采购决策、塑造公众信任并影响政策制定。然而,基准测试的一个关键假设是,通过API测得的模型性能能真实反映部署系统的行为。本文通过审计ChatGPT、Claude和Gemini在七个系统和九个基准上的表现,涵盖通用能力、社会偏见和迎合性,发现API与界面间存在系统性差异。平均而言,API评估的准确率比界面高3.4个百分点,测试-重测一致性高2.1个百分点。对于ChatGPT,API与界面之间的性能差距甚至超过GPT 5.3与GPT 5.4之间的差异。换言之,仅改变访问方式导致的性能下降,可能等同于更换一个完整模型版本。我们进一步测试了通过调整系统提示、采样参数和推理设置能否复现界面行为,结果显示部分情况可调节,但无法可靠消除差距。研究揭示了‘上下文有效性差距’:通过API获得的测量结果未必能推广到实际部署界面,使API评估难以作为部署系统的真实代理。
原文摘要 · Abstract (English)
Benchmark scores are a central currency in model releases: they inform purchasing decisions, shape public trust, and influence policy. Yet, a key assumption underlying benchmark scores is that the model performance measured through APIs faithfully reflects the behavior of deployed systems. We challenge this assumption by auditing ChatGPT, Claude, and Gemini across seven systems and nine benchmarks spanning general capability, social bias, and sycophancy. We find systematic API--interface differences in both accuracy and consistency. On average, API evaluations score 3.4 percentage points higher in accuracy and 2.1 percentage points higher in test--retest agreement than corresponding interface evaluations. For ChatGPT, the performance difference between API and interface access exceeds the API-only difference between GPT 5.3 and GPT 5.4. Put differently, switching access surfaces can degrade performance as much as downgrading a full model generation. We further test whether exposed API controls can reproduce interface behavior by varying system prompts, sampling parameters, and reasoning settings. These controls shift behavior in some cases but do not reliably eliminate the gap. Our findings document a context-validity gap: measurements obtained through APIs do not necessarily generalize to corresponding deployed interfaces, complicating the use of API evaluations as proxies for deployed systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。