评估大模型时,如何确保测试指标真正反映安全与鲁棒性?
Measuring what Matters: Construct Validity in Large Language Model Benchmarks

- 29位专家系统审查445个主流评测集,发现多数评测设计存在缺陷
- 多数评测指标与实际安全、鲁棒性等抽象能力关联不足,结论不可靠
- 提出8条可操作建议,指导研究者构建更有效的评测体系
评估大语言模型(LLMs)对于衡量其能力并提前发现安全或鲁棒性问题至关重要。可靠测量如‘安全’和‘鲁棒性’等抽象复杂现象,需要强建构效度,即测评指标真正代表所关注的特性。我们组织了29位专家,系统审查了自然语言处理与机器学习顶级会议中的445个LLM评测。分析发现,这些评测在测量对象、任务设计和评分机制上存在普遍问题,导致其结论的有效性受损。为解决这些问题,本文提出八项关键建议,并提供详细可操作的指导,帮助研究人员和实践者改进大语言模型评测的设计。
原文摘要 · Abstract (English)
Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as 'safety' and 'robustness' requires strong construct validity, that is, having measures that represent what matters to the phenomenon. With a team of 29 expert reviewers, we conduct a systematic review of 445 LLM benchmarks from leading conferences in natural language processing and machine learning. Across the reviewed articles, we find patterns related to the measured phenomena, tasks, and scoring metrics which undermine the validity of the resulting claims. To address these shortcomings, we provide eight key recommendations and detailed actionable guidance to researchers and practitioners in developing LLM benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。