arXiv:2502.03461cs.LGcs.CL2025-02被引 53

现有大模型评测存在标签错误,影响可靠性判断,作者提出更精准的铂金评测集。

Do Large Language Model Benchmarks Test Reliability?

  • 构建铂金评测集,通过修正15个主流基准的数据标注减少误差。
  • 前沿大模型在小学数学应用题上仍频繁出错,暴露未被发现的系统性缺陷。
  • 适合关注模型真实可靠性、评测体系设计的研究者和开发者。

部署大型语言模型时,不仅需评估其能力,还需确保其可靠性。尽管已有大量基准用于追踪模型能力进步,但对可靠性的测量却长期缺失。本文研究现有基准在量化模型可靠性方面的有效性,发现普遍存在的标签错误会掩盖模型的持续失败与不可靠行为。为此,我们提出‘铂金基准’概念——通过精心校准以最小化标签错误和歧义。作为初步尝试,我们修订了15个主流基准中的示例。在这些铂金基准上评估多种模型,发现前沿大模型在基础数学应用题等简单任务中仍频繁出错。进一步分析揭示了模型在特定类型问题上持续表现不佳的新型模式。代码已开源:https://github.com/MadryLab/platinum-benchmarks。

原文摘要 · Abstract (English)

When deploying large language models (LLMs), it is important to ensure that these models are not only capable, but also reliable. Many benchmarks have been created to track LLMs' growing capabilities, however there has been no similar focus on measuring their reliability. To understand the potential ramifications of this gap, we investigate how well current benchmarks quantify model reliability. We find that pervasive label errors can compromise these evaluations, obscuring lingering model failures and hiding unreliable behavior. Motivated by this gap in the evaluation of reliability, we then propose the concept of so-called platinum benchmarks, i.e., benchmarks carefully curated to minimize label errors and ambiguity. As a first attempt at constructing such benchmarks, we revise examples from fifteen existing popular benchmarks. We evaluate a wide range of models on these platinum benchmarks and find that, indeed, frontier LLMs still exhibit failures on simple tasks such as elementary-level math word problems. Analyzing these failures further reveals previously unidentified patterns of problems on which frontier models consistently struggle. We provide code at https://github.com/MadryLab/platinum-benchmarks

大模型评测可靠性铂金基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。