arXiv:2501.04227cs.HCcs.AI2025-01EMNLP被引 478

用大模型自动完成科研全流程,省时省力还省钱。

Agent Laboratory: Using LLM Agents as Research Assistants

  • 大模型自主开展文献调研、实验和报告撰写。
  • 生成代码性能达顶尖水平,成本比旧方法低84%。
  • 人参与反馈能显著提升研究质量,适合想提速的科研者。

科学发现历来耗时耗资,从构想到成果需投入大量人力物力。为加速科研进程、降低研究成本并提升质量,我们提出Agent Laboratory——一个基于大语言模型的自主科研框架,可完整执行研究流程。该框架接收人类研究设想后,依次经历文献综述、实验与报告撰写三个阶段,最终输出代码仓库与研究报告,并支持用户在各阶段提供反馈与指导。我们部署了多种前沿大模型进行测试,邀请多位研究人员参与调查,通过提供反馈引导研究并评估最终论文质量。结果显示:(1) o1-preview 驱动的 Agent Laboratory 产出效果最佳;(2) 自动生成的机器学习代码达到现有方法的领先水平;(3) 人在各阶段提供反馈显著提升整体研究质量;(4) 相较于以往自主研究方法,研究支出降低84%。我们期望该框架让研究人员将精力聚焦于创造性构思,而非重复性编码与写作,从而加速科学发现。

原文摘要 · Abstract (English)

Historically, scientific discovery has been a lengthy and costly process, demanding substantial time and resources from initial conception to final results. To accelerate scientific discovery, reduce research costs, and improve research quality, we introduce Agent Laboratory, an autonomous LLM-based framework capable of completing the entire research process. This framework accepts a human-provided research idea and progresses through three stages--literature review, experimentation, and report writing to produce comprehensive research outputs, including a code repository and a research report, while enabling users to provide feedback and guidance at each stage. We deploy Agent Laboratory with various state-of-the-art LLMs and invite multiple researchers to assess its quality by participating in a survey, providing human feedback to guide the research process, and then evaluate the final paper. We found that: (1) Agent Laboratory driven by o1-preview generates the best research outcomes; (2) The generated machine learning code is able to achieve state-of-the-art performance compared to existing methods; (3) Human involvement, providing feedback at each stage, significantly improves the overall quality of research; (4) Agent Laboratory significantly reduces research expenses, achieving an 84% decrease compared to previous autonomous research methods. We hope Agent Laboratory enables researchers to allocate more effort toward creative ideation rather than low-level coding and writing, ultimately accelerating scientific discovery.

AI科研自动化大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。