首个评估大模型设计科学消融实验能力的基准,发现模型表现远逊于人类。
AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research
- 构建1500个来自807篇NLP论文的专家标注数据集,评估模型设计消融实验能力。
- 主流大模型在重要性、忠实性和合理性上显著落后于人类专家。
- 提出AbGen-Eval,揭示现有自动化评估方法不可靠,适合研究评估系统改进。
我们提出AbGen,首个用于评估大语言模型在科学研宄中设计消融实验能力的基准。AbGen包含1500个由专家标注的示例,源自807篇NLP论文。在此基准中,模型需根据给定研究背景,为特定模块或过程生成详细的消融实验设计方案。对DeepSeek-R1-0528和o4-mini等领先模型的评估显示,其在重要性、忠实性和合理性方面与人类专家存在显著差距。此外,我们证明当前自动化评估方法不可靠,其结果与人工评估存在显著差异。为此,我们开发了AbGen-Eval,一个元评估基准,用于评估常用自动化评估系统在本任务中的可靠性。我们对多种LLM-as-Judge系统进行测试,为未来开发更有效、可靠的基于大模型的复杂科学任务评估系统提供洞见。
原文摘要 · Abstract (English)
We introduce AbGen, the first benchmark designed to evaluate the capabilities of LLMs in designing ablation studies for scientific research. AbGen consists of 1,500 expert-annotated examples derived from 807 NLP papers. In this benchmark, LLMs are tasked with generating detailed ablation study designs for a specified module or process based on the given research context. Our evaluation of leading LLMs, such as DeepSeek-R1-0528 and o4-mini, highlights a significant performance gap between these models and human experts in terms of the importance, faithfulness, and soundness of the ablation study designs. Moreover, we demonstrate that current automated evaluation methods are not reliable for our task, as they show a significant discrepancy when compared to human assessment. To better investigate this, we develop AbGen-Eval, a meta-evaluation benchmark designed to assess the reliability of commonly used automated evaluation systems in measuring LLM performance on our task. We investigate various LLM-as-Judge systems on AbGen-Eval, providing insights for future research on developing more effective and reliable LLM-based evaluation systems for complex scientific tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。