测试发现同一模型在不同接口和搜索条件下表现差异大,仅看准确率会掩盖安全风险。
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

- 对比API与聊天界面、有无网络搜索,多轮测试分析行为差异
- 开启搜索后准确率下降最多达8个百分点,甚至反转性能趋势
- 响应不一致率达21%,引用来源和回避回答行为也因接口而异
大型语言模型(LLM)评估常被用于支持模型安全性、可靠性及部署就绪性的声明。然而,多数评估仅依赖单一访问模态(如API)、单次运行每条提示,并以准确率为唯一指标,未考虑实际部署中可能影响模型行为的条件,如网络搜索。本文针对最广泛应用的LLM之一,比较ChatGPT的聊天界面与OpenAI API两种模态,在启用或禁用网络搜索的情况下进行评估。采用来自BBQ和SafetyBench两个主流基准的401个分层样本提示,每条提示执行三次重复运行,共收集4,812条响应。除标准性能指标外,还评估了响应一致性、文本相似度、引用依据性及回避行为等维度。结果显示:在禁用搜索时,聊天界面准确率低于API;启用搜索后,准确率最高下降8个百分点,且在某一基准上甚至逆转了模态间的性能趋势。同一提示的多次运行中,高达21%的提示出现不一致响应。两种模态引用的来源不同,回避行为也存在不一致。这些结果表明,即使在同一模型家族内,仅报告简单准确率会掩盖与安全评估相关的显著行为变异。我们主张,应系统性地将模态、多轮一致性、搜索条件及响应层面行为纳入人工智能安全评估,以更真实反映实际部署中的系统表现。
原文摘要 · Abstract (English)
Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet most evaluations rely on a single access modality (model APIs), perform a single run per prompt, and report accuracy as the primary outcome metric, without accounting for conditions such as web search that may have effects on model behavior in deployment. We audit these assumptions for one of the most widely-used LLMs, comparing two modalities, ChatGPT's chat UI and OpenAI's API, with and without web search enabled. We use a stratified total sample of 401 prompts from two popular benchmarks, BBQ and SafetyBench, collecting 4,812 total responses across three repeated runs per prompt. Beyond standard performance measures, we evaluate model output dimensions including response consistency, response text similarity, citation grounding, and abstention behavior. For instance, chat UI responses were less accurate than API responses on both benchmarks with search disabled. Enabling web search reduced accuracy by up to 8 percentage points, and even reversed the direction of modality performance trends for one benchmark. Repeated runs of the same prompt produced inconsistent responses in up to 21\% of prompts. The two modalities also grounded answers in different citations, and abstention behavior was also inconsistent across both modalities. These results illustrate that, even within a model family, reporting only simple accuracy metrics can obscure important forms of model behavioral variation relevant to AI safety assessments. We argue that AI safety evaluations should systematically account for modality, multi-run consistency, search conditions, and response-level behaviors to better reflect how deployed AI systems behave in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。