arXiv:2503.18102cs.AIcs.CL2025-03被引 63

让大模型科研代理在共享平台协作,持续迭代提升研究效率。

AgentRxiv: Towards Collaborative Autonomous Research

  • 构建共享预印本服务器,支持代理间上传下载报告实现协同研究。
  • 有历史记录的代理比孤立代理在MATH-500上提升11.4%性能。
  • 多代理协作使整体准确率提升13.7%,适用于跨领域任务优化。

科学发现通常是数百名科学家逐步协作的结果,而非单一灵感迸发。现有智能体工作流虽能自主生成研究,但彼此孤立,无法持续改进前序成果。为此,我们提出AgentRxiv框架,允许大语言模型代理实验室将研究报告上传至共享预印本服务器,实现信息共享与迭代合作。实验中,具备历史研究访问能力的代理在MATH-500任务上相比孤立代理取得11.4%的相对性能提升;多代理通过该平台协作时,在相同测试集上整体准确率提升13.7%。最佳策略在其他领域基准上也表现出平均3.3%的泛化增益。结果表明,自主智能体有望与人类共同设计未来人工智能系统。我们希望AgentRxiv推动代理间科研协作,加速科学发现进程。

原文摘要 · Abstract (English)

Progress in scientific discovery is rarely the result of a single "Eureka" moment, but is rather the product of hundreds of scientists incrementally working together toward a common goal. While existing agent workflows are capable of producing research autonomously, they do so in isolation, without the ability to continuously improve upon prior research results. To address these challenges, we introduce AgentRxiv-a framework that lets LLM agent laboratories upload and retrieve reports from a shared preprint server in order to collaborate, share insights, and iteratively build on each other's research. We task agent laboratories to develop new reasoning and prompting techniques and find that agents with access to their prior research achieve higher performance improvements compared to agents operating in isolation (11.4% relative improvement over baseline on MATH-500). We find that the best performing strategy generalizes to benchmarks in other domains (improving on average by 3.3%). Multiple agent laboratories sharing research through AgentRxiv are able to work together towards a common goal, progressing more rapidly than isolated laboratories, achieving higher overall accuracy (13.7% relative improvement over baseline on MATH-500). These findings suggest that autonomous agents may play a role in designing future AI systems alongside humans. We hope that AgentRxiv allows agents to collaborate toward research goals and enables researchers to accelerate discovery.

智能体协作研究自主科研

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。