用双层协作机制自动规划并执行科研方案,提升生成论文的质量与可行性。
Evolving and Executing Research Plans via Double-Loop Multi-Agent Collaboration
- 双层多智能体架构:导师组演化研究计划,学生组动态执行并调整
- 在ACLAward和Laboratory数据集上达到最先进水平,显著优于基线方法
- 适合自动化科研、AI for Science领域研究人员参考
自动化端到端科学研究面临根本挑战:既要生成新颖且合理的高层研究计划,又要在动态不确定环境下正确执行。为此,我们提出双层多智能体(DLMA)框架,实现研究问题的全自动求解。领导层由教授智能体组成,通过参与、改进与整合会议,利用进化算法迭代生成并优化研究提案,高效探索解决方案空间。执行层由博士生智能体构成,通过事前与事后会议动态调整执行计划,确保每一步(如撰写、编程)均获得上下文与外部观测的支持。在ACLAward和Laboratory等基准上的实验表明,DLMA生成的论文在自动评估中取得最先进成绩,显著优于强基线。消融实验证实两层均至关重要:演化层驱动新颖性,执行层保障严谨性。
原文摘要 · Abstract (English)
Automating the end-to-end scientific research process poses a fundamental challenge: it requires both evolving high-level plans that are novel and sound, and executing these plans correctly amidst dynamic and uncertain conditions. To address this bilevel challenge, we propose a novel Double-Loop Multi-Agent (DLMA) framework to solve the given research problem automatically. The leader loop, composed of professor agents, is responsible for evolving research plans. It employs an evolutionary algorithm through involvement, improvement, and integration meetings to iteratively generate and refine a pool of research proposals, exploring the solution space effectively. The follower loop, composed of doctoral student agents, is responsible for executing the best-evolved plan. It dynamically adjusts the plan during implementation via pre-hoc and post-hoc meetings, ensuring each step (e.g., drafting, coding) is well-supported by contextual and external observations. Extensive experiments on benchmarks like ACLAward and Laboratory show that DLMA generates research papers that achieve state-of-the-art scores in automated evaluation, significantly outperforming strong baselines. Ablation studies confirm the critical roles of both loops, with evolution driving novelty and execution ensuring soundness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。