用分层搜索与自验证提升视觉推理的准确性
Visual Attention Reasoning via Hierarchical Search and Self-Verification
- 将线性思维改为可回溯的树状搜索,增强逻辑纠错能力
- 通过显式框选和组合奖励机制,使答案有可追溯的视觉证据
- 在复杂推理和安全测试中显著优于现有方法
多模态大语言模型常因依赖脆弱的线性推理和弱视觉定位而产生幻觉。我们提出视觉注意力推理(VAR),一种强化学习框架,将推理重构为具有自验证能力的分层搜索。VAR通过新型奖励函数(结合几何精度与语义充分性)引导生成显式边界框,实现可追溯的视觉证据锚定。同时,以树搜索策略替代线性思维链,支持回溯修正逻辑错误。理论分析证明了该框架的可靠性,大量实验表明,VAR在复杂幻觉与安全基准上显著优于当前最先进方法。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) frequently hallucinate due to their reliance on fragile, linear reasoning and weak visual grounding. We propose Visual Attention Reasoning (VAR), a reinforcement learning framework that reformulates reasoning as a hierarchical search with self-verification. VAR enforces traceable evidence grounding by generating explicit bounding boxes, guided by a novel reward function combining geometric precision and semantic sufficiency. Furthermore, it replaces linear Chain-of-Thought with a tree-search policy capable of backtracking to correct logical errors. Theoretical analysis validates the framework's reliability, and extensive experiments demonstrate that VAR significantly outperforms state-of-the-art methods on complex hallucination and safety benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。