用多智能体协作修正视觉误读,提升数学图文推理准确率
M$^3$-ACE: Rectifying Visual Perception in Multimodal Math Reasoning via Multi-Agentic Context Engineering
- 拆分感知与推理,通过多智能体共享视觉证据列表协同纠错
- 在MathVision上达89.1分新高,其他数据集也稳定提升
- 适合需要精准视觉理解的多模态推理任务研究者
多模态大模型在视觉数学推理中表现虽佳,但常受限于视觉感知不准这一未被充分关注的瓶颈。系统分析发现,错误主要源于视觉信息提取不完整或错误,而非推理能力不足;且模型对初始感知过度自信,传统提示工程、多轮自省等方法难以有效纠正。为此,我们提出M3-ACE——一种多智能体上下文工程框架,通过动态维护以视觉证据列表为中心的共享上下文,解耦感知与推理。多个智能体协同贡献互补观察,暴露不一致并恢复缺失信息。为支持稳定多轮协作,引入两个轻量工具:摘要工具将不同智能体的证据整理为一致、互补、冲突三类;精炼工具过滤不可靠样本并引导迭代修正。大量实验表明,M3-ACE显著提升跨多个基准的视觉数学推理性能,在MathVision上达到89.1的新纪录,并在MathVista和MathVerse等数据集上实现持续改进。结果凸显了以感知为中心的多智能体协作对推进多模态推理系统的重要性。
原文摘要 · Abstract (English)
Multimodal large language models have recently shown promising progress in visual mathematical reasoning. However, their performance is often limited by a critical yet underexplored bottleneck: inaccurate visual perception. Through systematic analysis, we find that the most failures originate from incorrect or incomplete visual evidence extraction rather than deficiencies in reasoning capability. Moreover, models tend to remain overly confident in their initial perceptions, making standard strategies such as prompt engineering, multi-round self-reflection, or posterior guidance insufficient to reliably correct errors. To address this limitation, we propose M3-ACE, a multi-agentic context engineering framework designed to rectify visual perception in multimodal math reasoning. Instead of directly aggregating final answers, our approach decouples perception and reasoning by dynamically maintaining a shared context centered on visual evidence lists. Multiple agents collaboratively contribute complementary observations, enabling the system to expose inconsistencies and recover missing perceptual information. To support stable multi-turn collaboration, we further introduce two lightweight tools: a Summary Tool that organizes evidence from different agents into consistent, complementary, and conflicting components, and a Refine Tool that filters unreliable samples and guides iterative correction. Extensive experiments demonstrate that M3-ACE substantially improves visual mathematical reasoning performance across multiple benchmarks. Our method establishes new state-of-the-art results 89.1 on the MathVision benchmark and achieves consistent improvements on other related datasets, including MathVista and MathVerse. These results highlight the importance of perception-centric multi-agent collaboration for advancing multimodal reasoning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。