用真人专家对决方式评估大模型科研创意,发现部分框架能提升想法质量。
Ideation Arena: Evaluating LLM Generated Research Ideas with Battle-style Human Expert Assessment

- 通过双盲对决让专家对比大模型生成的科研想法
- 105位研究者完成超6000次评估,构建出计算机科学创意偏好排行榜
- 首次建立自动化评价基准,现有大模型评分仅72.56%准确
评估大模型生成的科研想法很困难,因其科学价值无法仅靠客观标准判断,也无统一参考答案。为此,我们提出Ideation Arena——一种基于对决形式的人类专家评估平台。该平台对14个前沿大模型及基于2个基础模型构建的5种研究代理架构生成的想法进行评估。为确保起点一致,所有模型与代理均在由研究人员熟悉的论文构建的共享文献背景下运行。我们收集了来自105位活跃计算机科学研究员的超过6000次双盲两两比较,并基于封闭式语境协议构建了提案阶段专家偏好的Elo评分排行榜。通过评分者一致性与鲁棒性分析验证,榜单在标注者构成和领域覆盖变化下依然稳定。结果表明,代理框架表现差异显著,部分优于其基础模型,部分则无提升甚至更差。我们进一步构建了Ideation Arena Eval,用于评估自动化评价器是否匹配人类偏好。实验显示当前大模型裁判仍无法可靠复现专家意见,最佳裁判在整体质量上仅达到72.56% Soft Accuracy。代码、数据与排行榜已开源。
原文摘要 · Abstract (English)
Evaluating research ideas generated by LLMs is difficult because their scientific value cannot be fully determined by objective criteria, and no single reference answer specifies what counts as a good idea. To address this challenge, we introduce Ideation Arena, a battle style platform that evaluates research ideas through pairwise human assessment. Ideation Arena evaluates ideas generated by 14 frontier LLMs and 5 research agent architectures built on 2 base models. To ensure a common starting point, Ideation Arena builds shared literature contexts from papers familiar to the participating researchers and provides the same contexts to all LLMs and agents. We collect over 6,000 double blind pairwise comparisons from 105 active computer science researchers and construct an Elo rating leaderboard of proposal-stage expert preferences in computer science under a shared closed-context protocol. We validate the rankings through interrater agreement and robustness analyses, showing that the leaderboard remains stable under changes in annotator composition and domain coverage. Our results show substantial variation in agent effectiveness, with some frameworks improving ideation quality over their backbones and others offering little benefit or even underperforming their base models. We further construct Ideation Arena Eval, a benchmark for assessing whether automated evaluators align with human preferences in research ideation. Experiments with current LLM judges show that they still cannot reliably reproduce expert preferences, with the best judge reaching 72.56% Soft Accuracy on Overall Quality. Our code, data, and leaderboards are available at https://github.com/foss12138/Research-Ideation-Arena.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。