评测大模型能否设计高质量实验,发现其配置能力普遍不足。
Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design

- 构建涵盖19个领域的300篇顶会论文的评估基准SCOPE
- 多数大模型无法直接生成高质量实验方案,低层配置是主要瓶颈
- 提出OptED框架,通过分阶段+工具增强提升实验设计质量
AI for Research(AI4Research)利用AI自动化并改进科研流程。尽管实验设计是研究过程的关键环节,但以往工作主要关注代码实现与执行,忽视了该阶段的重要性,且缺乏评估AI系统化实验设计能力的基准。为填补这一空白,我们提出了SCOPE——一个基于300篇顶级会议(如ICML、NeurIPS、ICLR)最新论文、覆盖19个研究领域的科学综合规划评估基准,从两个维度评估大模型:高层规划完整性(主实验、消融实验、分析实验)和底层配置准确性与合理性(数据集、基线、指标)。基准测试揭示三个发现:(1)大多数大模型无法直接设计高质量实验;(2)所有大模型在低层配置上存在性能瓶颈;(3)搜索模式无法提升设计质量。为此,我们提出OptED——一种新型智能体工作流,通过阶段隔离、工具增强和规则约束,有效缓解配置瓶颈,提升大模型实验规划能力。
原文摘要 · Abstract (English)
AI for Research (AI4Research) leverages AI to automate and improve scientific workflows. While experimental design is a critical stage of the research process, prior work has focused primarily on code implementation and execution, overlooking the importance of this stage, and no benchmark exists to evaluate AI's ability to conduct systematic experiment design. To bridge this gap, we propose SCOPE, a Scientific COmprehensive Planning Evaluation Benchmark constructed from 300 high-quality latest papers across 19 research domains from top-tier venues (e.g., ICML, NeurIPS, and ICLR),evaluating LLMs on two dimensions: High-Level planning completeness (main, ablation, and analysis experiments) and Low-Level configuration accuracy and rationality (datasets, baselines, and metrics). Benchmarking reveals three findings: (1) most LLMs cannot directly design high-quality experiments; (2) all LLMs exhibit a performance bottleneck in low-level configuration; and (3) search mode does not improve design quality. Furthermore, to address these challenges, we propose OptED, a novel agentic workflow to optimize LLM-based experimental design, that enhances LLM-based experimental planning through stage isolation, tool augmentation, and rule-based constraints, effectively alleviating the configuration bottleneck.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。