让AI看图反思,避免幻觉,提升视觉推理准确性。
Bridging Modality Disconnect in Self-Reflection via Closed-Loop Visually Grounded Verification
- 构建闭环反思机制,逐轮验证图像区域并修正答案。
- 在多个基准上准确率提升,幻觉现象显著减少。
- 适合需要高可靠性视觉推理的场景,如医疗、自动驾驶。
在视觉语言模型时代,增强多模态推理能力仍是关键挑战,尤其在处理模糊或复杂的视觉输入时,初始推断常导致幻觉或逻辑错误。现有模型虽能生成看似合理但缺乏图像依据的答案,即使被要求‘反思’,其修正仍可能脱离视觉证据。为此,我们提出MIRROR框架,即基于视觉区域的多模态迭代推理与反思。该框架采用闭合回路设计,包含草稿、批判、基于区域的验证和修订四个阶段,循环执行直至输出与图像证据一致。为支持训练,我们构建了ReflectV数据集,包含反射触发信号、基于区域的验证动作及以视觉证据为基础的答案修正。在通用视觉语言基准和代表性推理基准上的实验表明,MIRROR提升了正确性并减少了视觉幻觉,证明将反思训练为一种寻找证据、关注区域的验证过程,而非单纯的文本修正,具有显著价值。
原文摘要 · Abstract (English)
In the era of Vision-Language Models (VLMs), enhancing multimodal reasoning capabilities remains a critical challenge, particularly in handling ambiguous or complex visual inputs, where initial inferences often lead to hallucinations or logic errors. Existing VLMs often produce plausible yet ungrounded answers, and even when prompted to "reflect", their corrections may remain detached from the image evidence. To address this, we propose the MIRROR framework for Multimodal Iterative Reasoning via Reflection On visual Regions. By embedding visual reflection as a core mechanism, MIRROR is formulated as a closed-loop process comprising draft, critique, region-based verification, and revision, which are repeated until the output is visually grounded. To facilitate training of this model, we construct **ReflectV**, a visual reflective dataset for multi-turn supervision that explicitly contains reflection triggers, region-based verification actions, and answer revision grounded in visual evidence. Experiments on both general vision-language benchmarks and representative vision-language reasoning benchmarks show that MIRROR improves correctness and reduces visual hallucinations, demonstrating the value of training reflection as an evidence-seeking, region-aware verification process rather than a purely textual revision step.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。