arXiv:2505.18705cs.AI2025-05NeurIPS被引 86

AI-Researcher可自主完成从发想到写论文的全流程科研。

AI-Researcher: Autonomous Scientific Innovation

  • 用AI自动完成文献调研、假设生成到论文撰写全流程。
  • 在多个领域任务中实现接近人类水平的研究成果。
  • 适合想提升效率的科研人员或探索自动化研究的新方向。

大型语言模型(LLMs)在数学和编程中的强大推理能力,结合代理框架自动化复杂任务的能力,为加速科学创新带来了前所未有的机遇。本文提出AI-Researcher,一个完全自主的研究系统,重新定义了人工智能驱动的科学发现与评估方式。该框架无缝协调完整的科研流程——从文献综述、假设生成、算法实现到生成可发表的论文稿件——仅需极少的人工干预。为严格评估自主研究能力,我们构建了Scientist-Bench基准,涵盖多个AI研究领域的前沿论文,包含引导式创新与开放式探索任务。大量实验表明,AI-Researcher实现了显著的代码实现成功率,并产出质量接近人类水平的研究论文。这项工作为自主科学创新奠定了新基础,能够超越人类认知局限,系统性探索解决方案空间,辅助人类研究人员。

原文摘要 · Abstract (English)

The powerful reasoning capabilities of Large Language Models (LLMs) in mathematics and coding, combined with their ability to automate complex tasks through agentic frameworks, present unprecedented opportunities for accelerating scientific innovation. In this paper, we introduce AI-Researcher, a fully autonomous research system that transforms how AI-driven scientific discovery is conducted and evaluated. Our framework seamlessly orchestrates the complete research pipeline--from literature review and hypothesis generation to algorithm implementation and publication-ready manuscript preparation--with minimal human intervention. To rigorously assess autonomous research capabilities, we develop Scientist-Bench, a comprehensive benchmark comprising state-of-the-art papers across diverse AI research domains, featuring both guided innovation and open-ended exploration tasks. Through extensive experiments, we demonstrate that AI-Researcher achieves remarkable implementation success rates and produces research papers that approach human-level quality. This work establishes new foundations for autonomous scientific innovation that can complement human researchers by systematically exploring solution spaces beyond cognitive limitations.

自主科研AI研究智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。