用多智能体工作流提升化学多模态搜索,效果媲美GPT-4o
Agentic Mixture-of-Workflows for Multi-Modal Chemical Search
- 设计多智能体工作流,融合不同检索增强生成策略
- 在小分子、聚合物、反应与核磁谱图检索中表现接近GPT-4o
- 支持可解释性评估,适合科研自动化与AI模型对比研究
庞大的材料设计空间需要创新策略来整合多学科科学知识并优化材料发现。尽管大语言模型(LLMs)在多个领域展现出出色的推理与自动化能力,但其在材料科学中的应用仍受限于缺乏基准标准和实际实施框架。为此,我们提出自校正检索增强生成的多工作流混合(CRAG-MoW)——一种新型范式,通过开源LLM协调多种代理工作流,采用不同的CRAG策略。与以往方法不同,CRAG-MoW由编排代理合成多样输出,实现对同一问题域中多个LLM的直接评估。我们在小分子、聚合物、化学反应及多模态核磁共振(NMR)光谱检索任务上进行了基准测试。结果表明,CRAG-MoW性能与GPT-4o相当,且在对比评估中更受青睐,凸显结构化检索与多代理合成的优势。该方法揭示了不同数据类型下的性能差异,提供了一种可扩展、可解释且基于基准的AI架构优化路径,对填补科学应用中LLM与自主智能体的基准空白具有重要意义。
原文摘要 · Abstract (English)
The vast and complex materials design space demands innovative strategies to integrate multidisciplinary scientific knowledge and optimize materials discovery. While large language models (LLMs) have demonstrated promising reasoning and automation capabilities across various domains, their application in materials science remains limited due to a lack of benchmarking standards and practical implementation frameworks. To address these challenges, we introduce Mixture-of-Workflows for Self-Corrective Retrieval-Augmented Generation (CRAG-MoW) - a novel paradigm that orchestrates multiple agentic workflows employing distinct CRAG strategies using open-source LLMs. Unlike prior approaches, CRAG-MoW synthesizes diverse outputs through an orchestration agent, enabling direct evaluation of multiple LLMs across the same problem domain. We benchmark CRAG-MoWs across small molecules, polymers, and chemical reactions, as well as multi-modal nuclear magnetic resonance (NMR) spectral retrieval. Our results demonstrate that CRAG-MoWs achieve performance comparable to GPT-4o while being preferred more frequently in comparative evaluations, highlighting the advantage of structured retrieval and multi-agent synthesis. By revealing performance variations across data types, CRAG-MoW provides a scalable, interpretable, and benchmark-driven approach to optimizing AI architectures for materials discovery. These insights are pivotal in addressing fundamental gaps in benchmarking LLMs and autonomous AI agents for scientific applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。