构建首个端到端科学自主研究评测基准,验证AI能否真正复现真实科研成果。
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research

- 基于40个真实论文任务,用专家制定的多模态评分标准评估复现能力
- 最强AI仅平均21.5分(满分100),多数系统难以复现实验协议与核心发现
- 适合关注AI科研自动化进展的研究者和评测框架开发者
AI编程代理在科学工作中日益普及,但其端到端自主研究能力仍难以验证。我们提出ResearchClawBench,一个涵盖10个科学领域40个任务的评测基准。每个任务均基于真实发表论文,提供相关文献与原始数据,并在评估时隐藏目标论文。专家设计的多模态评分体系将目标科研成果分解为加权指标,可在保证复现目标论文的同时允许新发现。我们在统一协议下评估了七种自主研究(auto-research)代理及十七个原生大语言模型(LLM),使用轻量级ResearchHarness框架。当前系统距离可靠复现仍有显著差距:最强代理Claude Code平均得分21.5,最强LLM Claude-Opus-4.7得分为20.7,整体前沿均值仅26.5。错误分析显示失败主要集中在实验协议不匹配、证据不符以及缺失科学核心。ResearchClawBench为衡量自主科学研究进展提供了可复现的评估前沿。
原文摘要 · Abstract (English)
AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify. We present ResearchClawBench, a benchmark for evaluating autonomous scientific research across 40 tasks from 10 scientific domains. Each task is grounded in a real published paper, provides related literature and raw data, and hides the target paper during evaluation. Expert-curated multimodal rubrics decompose the target scientific artifacts into weighted criteria, enabling evaluation of target-paper-level re-discovery while leaving room for new discovery. We evaluate seven autonomous research (auto-research) agents under a unified protocol and seventeen native LLMs through the lightweight ResearchHarness. Current systems remain far from reliable re-discovery: the strongest autonomous agent, Claude Code, averages 21.5, and the strongest ResearchHarness LLM, Claude-Opus-4.7, averages 20.7, with an LLM frontier mean of only 26.5. Error analysis shows that failures concentrate in experimental protocol mismatch, evidence mismatch, and missing scientific core. ResearchClawBench provides a reproducible evaluation frontier for measuring progress toward autonomous scientific research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。