用动态图结构替代线性推理,提升科学任务的准确性与鲁棒性。
TopoAgent: A Self-Evolving Topological Agent for Multimodal Scientific Reasoning

- 将复杂问题拆解为视觉关联的原子单元,构建有向无环图组织依赖关系
- 运行时自适应分裂瓶颈节点,突破工具能力边界,避免错误累积
- 在数学物理化学多领域测试中显著优于现有方法,适合复杂科学推理场景
尽管多模态大语言模型在通用任务中表现优异,但在严格科学推理方面仍面临挑战,主要源于单一、线性的规划模式。此类设计常导致视觉-语义错位、长上下文幻觉以及固定任务粒度下的脆弱执行。我们提出TopoAgent,一种自演化拓扑框架,以动态、状态隔离的图演化替代线性轨迹。TopoAgent首先通过前端分解器将复杂查询分解为视觉锚定的原子单元,并基于依赖关系构建有向无环图(DAG),实现严格上下文隔离,防止历史噪声干扰推理引擎。此外,引入自适应原子裂解机制,在运行时当工具能力边界被突破时,动态将瓶颈节点拆分为更细粒度的子原子。在数学、物理、化学基准上的大量实验表明,TopoAgent显著优于当前最先进的线性智能体框架,提供了一种鲁棒、抗噪声且具备自修正能力的自主科学推理范式。
原文摘要 · Abstract (English)
While Multimodal Large Language Models (MLLMs) excel in general tasks, rigorous scientific reasoning remains challenging due to the limitations of monolithic, linear planning. Such sequential designs often suffer from visual-semantic misalignment, long-context hallucinations, and brittle execution under fixed task granularity. We propose TopoAgent, a self-evolving topological framework that replaces linear trajectories with dynamic, state-isolated graph evolution. TopoAgent first employs a front-end decomposer to fracture complex queries into visually-grounded atoms. These atoms are organized into a Directed Acyclic Graph (DAG) based on their dependencies, enabling strict context isolation to shield the reasoning engine from irrelevant historical noise. Furthermore, we introduce adaptive atomic fission, which dynamically splits bottleneck nodes into finer-grained sub-atoms at runtime when tool capability boundaries are exceeded. Extensive experiments across mathematics, physics, and chemistry benchmarks demonstrate that TopoAgent significantly outperforms state-of-the-art linear agent frameworks, providing a robust, noise-resistant, and self-correcting paradigm for autonomous scientific reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。