让多模态大模型能自我纠错,提升工具辅助推理的可靠性。
Think, Look, and Revise: Inconsistency-Aware Visual Self-Correction in MLLMs

- 引入反思性训练数据,让模型学会验证工具输出
- 通过强化学习奖励机制减少空间错位,提升判断准确率
- 适合需要高可靠视觉推理的应用场景
工具增强的多模态推理将外部工具(如目标检测、深度估计)融入多模态大语言模型(MLLMs),以解决复杂视觉任务中的感知瓶颈。然而,现有方法极少验证工具输出,限制了其检测和恢复工具失效的能力。我们提出 ReVISE 框架,为 MLLMs 提供验证与动态错误恢复能力。ReVISE 引入(1)一个精心构建的训练数据集,监督模型的反思行为,使其能够验证工具生成的证据,在视觉不一致时重构问题,并在外部工具不可靠时回退到内在视觉基础;(2)基于强化学习的定向奖励机制,鼓励内部反思并惩罚空间错位。多个基准测试结果表明,该方法显著优于现有方法,凸显了在工具增强型多模态推理中错误检测与纠正的重要性。
原文摘要 · Abstract (English)
Tool-augmented multimodal reasoning integrates external tools (e.g., object detection, depth estimation) into multimodal large language models (MLLMs) to address perceptual bottlenecks in complex visual tasks. However, existing approaches rarely verify tool outputs, limiting their ability to detect and recover from tool failures. We propose ReVISE, a framework that equips MLLMs with verification and dynamic error recovery for tool-augmented reasoning. ReVISE introduces (1) a curated training dataset that supervises reflective behaviors, enabling models to validate tool-derived evidence, reformulate queries when visual mismatches arise, and fall back to intrinsic grounding when external tools are unreliable; and (2) a reinforcement learning based targeted rewards that encourage internal reflection and penalize spatial misalignment. Experiments on several benchmarks demonstrate consistent improvements over existing methods, highlighting the importance of error detection and correction in tool-augmented multimodal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。