arXiv:2410.12974cs.CL2024-10NeurIPS被引 11

为大模型评测提供标准化文档框架,提升选型透明度。

BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks

  • 设计统一文档模板,规范评测目标、方法与数据来源
  • 用户研究证实能显著降低选型难度与误用风险
  • 适合需要严谨评估大模型性能的研究者与开发者

大语言模型在多样化任务中表现强大,但不同模型在各领域能力差异显著。面对众多评测基准,用户难以甄别合适选项,易导致误用与误读,且需耗费大量精力筛选。为此,我们提出 exttt{BenchmarkCards}——一种直观且经验证的标准化文档框架,统一规范评测的目标、方法、数据来源与局限性等关键属性。通过面向评测创建者与使用者的用户研究,我们验证该框架可有效简化基准选择过程,增强评估透明度,支持更明智的大模型评价决策。数据与代码已公开于 https://github.com/SokolAnn/BenchmarkCards。

原文摘要 · Abstract (English)

Large language models (LLMs) are powerful tools capable of handling diverse tasks. Comparing and selecting appropriate LLMs for specific tasks requires systematic evaluation methods, as models exhibit varying capabilities across different domains. However, finding suitable benchmarks is difficult given the many available options. This complexity not only increases the risk of benchmark misuse and misinterpretation but also demands substantial effort from LLM users, seeking the most suitable benchmarks for their specific needs. To address these issues, we introduce \texttt{BenchmarkCards}, an intuitive and validated documentation framework that standardizes critical benchmark attributes such as objectives, methodologies, data sources, and limitations. Through user studies involving benchmark creators and users, we show that \texttt{BenchmarkCards} can simplify benchmark selection and enhance transparency, facilitating informed decision-making in evaluating LLMs. Data & Code: https://github.com/SokolAnn/BenchmarkCards

评测标准大模型文档框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。