arXiv:2510.21460cs.SEcs.CY2025-10NeurIPS被引 4

为大模型评测工具设计风险评估框架,帮用户判断评测结果是否可信。

Risk Management for Mitigating Benchmark Failure Modes: BenchRisk

  • 基于标准风险管理流程,系统分析26个评测集的潜在失效模式。
  • 识别出57种失效风险,提出196项缓解策略,构建可评分的BenchRisk体系。
  • 适合关注评测可靠性、想对比不同评测工具的研究者和部署者。

大型语言模型(LLM)评测影响模型使用决策(如“该模型在特定场景下是否安全可用?”)。然而,评测可能因偏差、方差、覆盖不足或使用者理解困难等失效模式而不可靠。本研究以美国国家标准与技术研究院(NIST)的风险管理流程为基础,迭代分析26个主流评测集,识别出57种潜在失效模式及196项对应缓解策略。这些策略可降低失效发生概率或严重性,形成“评测风险”评估框架,生成元评估指标BenchRisk。分数越高,表明用户得出错误或无支持结论的可能性越低。所有26个被测评测在至少一个维度(全面性、可理解性、一致性、正确性、持久性)上均显示显著风险,揭示了该领域亟待研究的方向。BenchRisk工作流支持评测间比较,作为开源工具,还促进风险与缓解措施的识别与共享。

原文摘要 · Abstract (English)

Large language model (LLM) benchmarks inform LLM use decisions (e.g., "is this LLM safe to deploy for my use case and context?"). However, benchmarks may be rendered unreliable by various failure modes that impact benchmark bias, variance, coverage, or people's capacity to understand benchmark evidence. Using the National Institute of Standards and Technology's risk management process as a foundation, this research iteratively analyzed 26 popular benchmarks, identifying 57 potential failure modes and 196 corresponding mitigation strategies. The mitigations reduce failure likelihood and/or severity, providing a frame for evaluating "benchmark risk," which is scored to provide a metaevaluation benchmark: BenchRisk. Higher scores indicate that benchmark users are less likely to reach an incorrect or unsupported conclusion about an LLM. All 26 scored benchmarks present significant risk within one or more of the five scored dimensions (comprehensiveness, intelligibility, consistency, correctness, and longevity), which points to important open research directions for the field of LLM benchmarking. The BenchRisk workflow allows for comparison between benchmarks; as an open-source tool, it also facilitates the identification and sharing of risks and their mitigations.

评测风险大模型风险管理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。