多智能体协作解析化学反应图,准确率超现有方法。
MACReD: A Multi-Agent Collaborative Reasoning Framework for Reaction Diagram Parsing

- 分层多智能体协同处理分子识别、箭头理解等任务
- 在硬匹配和软匹配下分别达75.2%与84.6%的F1分数
- 适合复杂反应图解析,尤其多步或树状结构
从科学文献中解析化学反应图面临布局异质、视觉元素交织及识别与推理融合困难等问题。现有视觉语言模型虽提升多模态理解能力,但在复杂图上仍难以保持空间连贯性并整合多维信息。为此,我们提出MACReD,一种分层多智能体框架,在统一的VLM引导架构中协调分子感知、箭头理解、文本提取与反应重构等专用智能体。规划与感知层采用细粒度检测应对视觉复杂性,推理层通过多图融合机制整合异构线索并强制化学一致的全局推理。在RxnScribe基准测试中,MACReD达到75.2%和84.6%的F1分数(硬匹配与软匹配),优于基线模型的69.1%和80.0%。结果表明,该方法对多样布局,包括多步和树状反应,均具强鲁棒性。
原文摘要 · Abstract (English)
Parsing chemical reaction diagrams from scientific literature is challenging due to heterogeneous layouts, intertwined visual elements, and the difficulty of integrating recognition and reasoning. Existing vision-language models advance multimodal understanding but still fail on complex diagrams, struggling to maintain spatial coherence and to integrate multidimensional information during reasoning. To address these issues, we propose MACReD, a hierarchical multi-agent framework that coordinates specialized agents for molecular perception, arrow understanding, text extraction, and reaction reconstruction within a unified VLM-guided architecture. The planning and perception layers use flexible, fine-grained detection to handle visual complexity, while the reasoning layer uses a multigraph fusion mechanism to integrate heterogeneous cues and enforce chemically consistent global reasoning. Experiments on the RxnScribe benchmark show that MACReD achieves state-of-the-art performance, with F1 scores of 75.2% and 84.6% under hard and soft match criteria, outperforming the RxnScribe baseline, which obtains 69.1% and 80.0%, respectively. These results demonstrate the robustness of MACReD across diverse diagram layouts, including multi-step and tree-structured reactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。