用真实论文任务测试AI研究代理,发现其能力与可靠性严重不匹配。
ResearchGym: Evaluating Language Model Agents on Real-World AI Research
- 基于5篇顶会论文构建可复现的科研环境,含39个子任务
- GPT-5代理仅在15次评估中成功超越基线1次,平均完成26.5%任务
- 揭示长周期失败模式,适合关注自主科研代理的研究者
我们提出ResearchGym,一个用于评估人工智能代理在端到端科研任务中表现的基准和执行环境。基于ICML、ICLR和ACL的五篇口头报告和亮点论文,我们保留原始数据集、评估框架和基线实现,但隐藏论文提出的具体方法,形成五个容器化任务环境,共包含39个子任务。代理需提出新假设、运行实验,并在论文指标上超越人类强基线。对搭载GPT-5的代理进行受控评估发现,其在15次评估中仅在1次成功超越基线(6.7%),平均仅完成26.5%的子任务。我们识别出重复出现的长周期失败模式,包括急躁、时间与资源管理差、对弱假设过度自信、难以协调并行实验,以及上下文长度限制。然而,在单次运行中,该代理仍超越了一项ICML 2025亮点任务的解决方案,表明前沿代理偶尔可达顶尖水平,但极不可靠。我们还评估了Claude Code(Opus-4.5)和Codex(GPT-5.2)等专有代理框架,均表现出类似的能力-可靠性差距。ResearchGym为闭环科研中自主代理的系统性评估与分析提供基础设施。
原文摘要 · Abstract (English)
We introduce ResearchGym, a benchmark and execution environment for evaluating AI agents on end-to-end research. To instantiate this, we repurpose five oral and spotlight papers from ICML, ICLR, and ACL. From each paper's repository, we preserve the datasets, evaluation harness, and baseline implementations but withhold the paper's proposed method. This results in five containerized task environments comprising 39 sub-tasks in total. Within each environment, agents must propose novel hypotheses, run experiments, and attempt to surpass strong human baselines on the paper's metrics. In a controlled evaluation of an agent powered by GPT-5, we observe a sharp capability--reliability gap. The agent improves over the provided baselines from the repository in just 1 of 15 evaluations (6.7%) by 11.5%, and completes only 26.5% of sub-tasks on average. We identify recurring long-horizon failure modes, including impatience, poor time and resource management, overconfidence in weak hypotheses, difficulty coordinating parallel experiments, and hard limits from context length. Yet in a single run, the agent surpasses the solution of an ICML 2025 Spotlight task, indicating that frontier agents can occasionally reach state-of-the-art performance, but do so unreliably. We additionally evaluate proprietary agent scaffolds including Claude Code (Opus-4.5) and Codex (GPT-5.2) which display a similar gap. ResearchGym provides infrastructure for systematic evaluation and analysis of autonomous agents on closed-loop research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。