评测AI智能体在真实大模型研究中的创新能力,挑战远超传统基准。
InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
- 构建20项全流程研究任务,要求生成可运行代码并评估结果质量。
- 前沿模型在复杂任务中表现不足,需超11小时才能达最佳性能。
- 适合研究自主科研智能体、长周期决策与代码生成的学者使用。
AI智能体可通过自动化假设生成、实验设计、编码、执行与分析加速科学发现,但现有基准仅测试简化场景下的单一技能。为此,我们提出InnovatorBench,一个用于真实、端到端评估智能体进行大语言模型(LLM)研究的基准-平台体系。该体系包含20个涵盖数据构建、过滤、增强、损失设计、奖励设计与架构构建的任务,要求生成可运行的成果,并评估正确性、性能、输出质量与不确定性。为支持智能体运行,我们开发ResearchGym研究环境,提供丰富的动作空间、分布式与长周期执行、异步监控及快照保存功能。同时,实现轻量级ReAct智能体,结合显式推理与可执行规划,采用Claude-4、GPT-5、GLM-4.5和Kimi-K2等前沿模型。实验表明,尽管前沿模型在代码驱动任务中表现良好,但在脆弱算法任务和长周期决策上仍存在缺陷,如急躁、资源管理差、过度依赖模板推理。此外,智能体需超过11小时才能在InnovatorBench上达到最优表现,凸显其难度,证明其具备成为下一代基于代码的研究基准的潜力。
原文摘要 · Abstract (English)
AI agents could accelerate scientific discovery by automating hypothesis formation, experiment design, coding, execution, and analysis, yet existing benchmarks probe narrow skills in simplified settings. To address this gap, we introduce InnovatorBench, a benchmark-platform pair for realistic, end-to-end assessment of agents performing Large Language Model (LLM) research. It comprises 20 tasks spanning Data Construction, Filtering, Augmentation, Loss Design, Reward Design, and Scaffold Construction, which require runnable artifacts and assessment of correctness, performance, output quality, and uncertainty. To support agent operation, we develop ResearchGym, a research environment offering rich action spaces, distributed and long-horizon execution, asynchronous monitoring, and snapshot saving. We also implement a lightweight ReAct agent that couples explicit reasoning with executable planning using frontier models such as Claude-4, GPT-5, GLM-4.5, and Kimi-K2. Our experiments demonstrate that while frontier models show promise in code-driven research tasks, they struggle with fragile algorithm-related tasks and long-horizon decision making, such as impatience, poor resource management, and overreliance on template-based reasoning. Furthermore, agents require over 11 hours to achieve their best performance on InnovatorBench, underscoring the benchmark's difficulty and showing the potential of InnovatorBench to be the next generation of code-based research benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。