arXiv:2506.22598cs.CL2025-06ACL被引 16

测试大模型智能体能否自主实现AI研究扩展,结果大多失败。

RExBench: Can coding agents autonomously implement AI research extensions?

论文配图:RExBench: Can coding agents autonomously implement AI research extensions?
图 1 · 摘自论文原文
  • 设计真实研究扩展任务,评估智能体自主实现能力
  • 12个智能体平均成功率仅33%,最高不足44%
  • 适合关注AI研究自动化与智能体能力边界的读者

基于大语言模型(LLMs)的智能体在自主执行复杂软件工程任务方面展现出潜力。近年来,也有进展推动智能体在机器学习与自然科学的研究流程中承担部分工作。我们认为,研究扩展及其实施是此类系统的关键能力,因此提出RExBench来评估该能力。RExBench是一个包含12篇论文真实扩展任务的基准,旨在检验新颖的研究假设。每个任务均为对已有论文和代码库的延伸,配有领域专家编写的指导说明。该基准对数据污染具有鲁棒性,并支持自动评估机制,通过执行智能体输出判断是否满足成功标准。我们使用该基准评估了12个基于aider和OpenHands框架实现的LLM智能体,发现所有智能体均无法自主完成多数扩展任务,最佳智能体成功率约为33%。尽管增加人工提示后成功率提升,但最高仍低于44%。这表明当前智能体仍难以在无大量人类指导的情况下处理真实的科研扩展任务。

原文摘要 · Abstract (English)

Agents based on Large Language Models (LLMs) have shown promise for performing sophisticated software engineering tasks autonomously. In addition, there has been progress towards developing agents that can perform parts of the research pipeline in machine learning and the natural sciences. We argue that research extension and its implementation is a critical capability for such systems, and introduce RExBench to support the evaluation of this capability. RExBench is a benchmark consisting of realistic extensions of 12 research papers that aim to investigate novel research hypotheses. Each task is set up as an extension to an existing research paper and codebase, accompanied by domain expert-written instructions. RExBench is robust to data contamination and supports an automatic evaluation infrastructure that executes agent outputs to determine whether the success criteria are met. We use this benchmark to evaluate 12 LLM agents implemented using two different frameworks, aider and OpenHands. We find that all agents fail to autonomously implement the majority of the extensions, with the best agent achieving around a 33% success rate. Although the success rate improves with additional human-written hints, the best performance under this setting remains below 44%. This indicates that current agents are still short of being able to handle realistic research extension tasks without substantial human guidance.

智能体研究自动化基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。