AI基准测试常被当作营销工具,实际决策中却未必可靠。
More than Marketing? On the Information Value of AI Benchmarks for Practitioners
- 通过访谈19人,发现基准测试主要用作模型性能相对对比信号。
- 学术界认为基准反映研究进展,但产品与政策领域多认为其不具实际参考性。
- 真正有用的基准需贴近真实场景、有领域知识支撑且透明可复现。
公开的AI基准测试结果被模型开发者广泛宣传,作为模型质量的指标,但在日益竞争的市场中,这些分数未必反映实际应用者关心的特性。本文基于对19位曾使用或拒绝使用基准测试人员的访谈分析发现,参与者普遍将基准测试视为模型间相对性能差异的信号,但其是否构成决策依据则因场景而异。在学术界,公开基准通常被视为衡量研究进展的有效方式;而在产品和政策领域,即使内部开发的基准也常被认为不足以支持实质性决策。受访者指出,不理想的基准往往因目标模糊、脱离真实应用场景而失效。因此,有效的基准应提供有意义的真实世界评估,融入领域专业知识,明确界定范围与目标,涵盖多样化的任务相关能力,具备足够挑战性以避免快速饱和,并考虑性能权衡而非依赖单一得分。此外,专有数据收集与防污染机制对确保结果可靠性至关重要。遵循这些标准,基准测试才能超越营销噱头,成为可靠的评估框架。
原文摘要 · Abstract (English)
Public AI benchmark results are widely broadcast by model developers as indicators of model quality within a growing and competitive market. However, these advertised scores do not necessarily reflect the traits of interest to those who will ultimately apply AI models. In this paper, we seek to understand if and how AI benchmarks are used to inform decision-making. Based on the analyses of interviews with 19 individuals who have used, or decided against using, benchmarks in their day-to-day work, we find that across these settings, participants use benchmarks as a signal of relative performance difference between models. However, whether this signal was considered a definitive sign of model superiority, sufficient for downstream decisions, varied. In academia, public benchmarks were generally viewed as suitable measures for capturing research progress. By contrast, in both product and policy, benchmarks -- even those developed internally for specific tasks -- were often found to be inadequate for informing substantive decisions. Of the benchmarks deemed unsatisfactory, respondents reported that their goals were neither well-defined nor reflective of real-world use. Based on the study results, we conclude that effective benchmarks should provide meaningful, real-world evaluations, incorporate domain expertise, and maintain transparency in scope and goals. They must capture diverse, task-relevant capabilities, be challenging enough to avoid quick saturation, and account for trade-offs in model performance rather than relying on a single score. Additionally, proprietary data collection and contamination prevention are critical for producing reliable and actionable results. By adhering to these criteria, benchmarks can move beyond mere marketing tricks into robust evaluative frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。