让视觉推理有据可循,模型不再‘假装看图’。
From Illusion to Intention: Visual Rationale Learning for Vision-Language Reasoning
- 将视觉操作视为核心推理步骤,而非可选工具。
- 在多个基准上达到顶尖性能,显著减少幻觉。
- 适合想提升视觉推理透明度的研究者和开发者。
近期视觉语言推理进展强调了以图像为思考基础的重要性,但现有框架将视觉操作视为可选工具,虽提升指标却使推理脱离真实视觉证据,导致‘看似看图’的假象。本文提出视觉理性学习(ViRL),将视觉操作重构为不可分割的推理单元,类似文本的思维链。ViRL采用端到端强化学习,包含:(1) 过程监督,使用真实推理路径;(2) 目标对齐,通过步骤级奖励引导;(3) 细粒度贡献分配,区分正确、冗余与错误动作。确保每一步都有效推动推理,实现‘因正确视觉理由而得正确答案’。仅通过端到端强化学习训练,ViRL在感知、幻觉、推理等多类任务上均达当前最优,建立了一种任务无关、过程可解释的可信视觉语言建模范式。
原文摘要 · Abstract (English)
Recent advances in vision-language reasoning underscore the importance of thinking with images, where models actively ground their reasoning in visual evidence. Yet, prevailing frameworks treat visual actions as optional tools, boosting metrics but leaving reasoning ungrounded and crops ineffective. This gap gives rise to the illusion of thinking with images: models seem visually grounded but rely on context-agnostic actions that neither refine perception nor guide reasoning toward correct answers. We address this problem by reframing visual actions as core reasoning primitives rather than optional tools, which we term visual rationalization, the visual analogue of textual Chain-of-Thought. Building on this insight, we propose Visual Rationale Learning (ViRL), an end-to-end paradigm that grounds training in the visual rationale itself. ViRL integrates (1) Process Supervision with ground-truth rationales, (2) Objective Alignment via step-level reward shaping, and (3) Fine-Grained Credit Assignment to distinguish correct, redundant, and erroneous actions. By ensuring each action contributes meaningfully to the reasoning chain, ViRL enables models to "get the right answer for the right visual reason". Trained purely with end-to-end RL, ViRL achieves state-of-the-art results across benchmarks spanning perception, hallucination, and reasoning. This work establishes visual rationalization as a task-agnostic, process-grounded paradigm for building transparent, verifiable, and trustworthy vision-language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。