arXiv:2506.18348cs.AI2025-06综述被引 5

让AI团队像真人一样讨论并互评,提升科研创意质量

Dynamic Knowledge Exchange and Dual-diversity Review: Concisely Unleashing the Potential of a Multi-Agent Research Team

  • AI研究员间动态交换知识,持续优化想法
  • 引入双多样性评审,生成更创新的科学假设
  • 在计算机与健康科学数据集上均超越现有系统

科学进步越来越依赖研究人员间的有效协作,而大语言模型(LLMs)才刚开始模拟这种动态。尽管近期基于LLM的科学家代理在自主科研发现中展现出潜力,但往往缺乏真实研究中必需的交互式推理与评估机制。我们提出IDVSCI(内部讨论与投票科学家),一个基于LLM的多智能体框架,包含两项关键创新:动态知识交换机制,支持智能体间迭代反馈;双多样性评审范式,模拟异质专家评估。这两者共同促进更深入的推理和更具创造性和影响力的科学构想生成。为评估方法的有效性与泛化能力,我们在两个数据集上进行实验:一个广泛使用的计算机科学基准,以及我们新引入的健康科学领域数据集。结果表明,IDVSCI在两个数据集上均持续表现最佳,优于AI Scientist和VIRSCI等现有系统。这些发现凸显了在基于LLM的自主研究中建模互动与同行评审动态的价值。

原文摘要 · Abstract (English)

Scientific progress increasingly relies on effective collaboration among researchers, a dynamic that large language models (LLMs) have only begun to emulate. While recent LLM-based scientist agents show promise in autonomous scientific discovery, they often lack the interactive reasoning and evaluation mechanisms essential to real-world research. We propose IDVSCI (Internal Discussion and Vote SCIentists), a multi-agent framework built on LLMs that incorporates two key innovations: a Dynamic Knowledge Exchange mechanism enabling iterative feedback among agents, and a Dual-Diversity Review paradigm that simulates heterogeneous expert evaluation. These components jointly promote deeper reasoning and the generation of more creative and impactful scientific ideas. To evaluate the effectiveness and generalizability of our approach, we conduct experiments on two datasets: a widely used benchmark in computer science and a new dataset we introduce in the health sciences domain. Results show that IDVSCI consistently achieves the best performance across both datasets, outperforming existing systems such as AI Scientist and VIRSCI. These findings highlight the value of modeling interaction and peer review dynamics in LLM-based autonomous research.

多智能体科研自动化知识交互创新生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。