arXiv:2603.18019cs.CLcs.AI2026-03

用检索工具揭示评测集是否真测了想测的能力

BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity

  • 构建检索系统,从20个评测套件中找与实际语言场景相关的测试项
  • 实验证明检索精准率高,能发现评测覆盖不全和排名不稳定问题
  • 适合关心评测可信度的研究者和模型开发者

语言模型评测是否真的测到了从业者想要的能力?现有高阶元数据过于粗略,无法反映评测的细节:一个名为‘诗歌’的评测可能从未考察俳句,而‘指令遵循’评测往往混合多种技能,难以判断真实意图。这种不透明使验证对齐实践目标变得繁琐,可能导致模型在未被测试的用户需求方面失败却仍被误认为有效。我们提出BenchBrowser,一个可从20个评测套件中检索出与自然语言使用场景相关评估项的检索器。通过人工研究验证,其检索精度高。BenchBrowser能生成证据,帮助从业者诊断内容效度低(能力维度覆盖不足)和聚合效度低(测量同一能力时排名不稳定)问题。该工具量化了实践意图与评测实际内容之间的关键差距。

原文摘要 · Abstract (English)

Do language model benchmarks actually measure what practitioners intend them to ? High-level metadata is too coarse to convey the granular reality of benchmarks: a "poetry" benchmark may never test for haikus, while "instruction-following" benchmarks will often test for an arbitrary mix of skills. This opacity makes verifying alignment with practitioner goals a laborious process, risking an illusion of competence even when models fail on untested facets of user interests. We introduce BenchBrowser, a retriever that surfaces evaluation items relevant to natural language use cases over 20 benchmark suites. Validated by a human study confirming high retrieval precision, BenchBrowser generates evidence to help practitioners diagnose low content validity (narrow coverage of a capability's facets) and low convergent validity (lack of stable rankings when measuring the same capability). BenchBrowser, thus, helps quantify a critical gap between practitioner intent and what benchmarks actually test.

评测验证语言模型效度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。