arXiv:2604.24339cs.CVcs.AI2026-04被引 1

通过视觉细节与反思机制,提升视觉语言模型的推理能力。

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection

论文配图:See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection
图 1 · 摘自论文原文
  • 引入低级视觉工具增强细粒度特征捕捉能力
  • 设计掩码反馈机制实现动态答案修正,准确率显著提升
  • 适合需要强视觉推理的科研与工业应用

近期视觉语言模型(VLM)在强化学习驱动下取得了进展,但仍存在忽视低级视觉信息和缺乏有效视觉反馈的问题。本文提出统一的多模态交错推理框架ForeSight,使模型能‘看更远’(利用低级视觉线索)并‘想更深’(通过视觉反馈反思)。首先,引入一组低级视觉工具,将关键视觉信息融入推理链,缓解对细粒度特征的忽略;其次,设计基于掩码的视觉反馈机制,让模型动态重审并更新答案。在强化学习驱动下,ForeSight自主决定工具调用与答案验证,以最终准确率为奖励信号。为评估性能,构建新数据集CG-SalBench(基于SalBench)。实验表明,ForeSight-7B在同规模模型中显著领先,甚至超越部分当前最先进的闭源模型。

原文摘要 · Abstract (English)

Recent advances in Vision-Language Models (VLMs) have benefited from Reinforcement Learning (RL) for enhanced reasoning. However, existing methods still face critical limitations, including the lack of low-level visual information and effective visual feedback. To address these problems, this paper proposes a unified multimodal interleaved reasoning framework \textbf{ForeSight}, which enables VLMs to \textbf{See Further} with low-level visual cues and \textbf{Think Deeper} with effective visual feedback. First, it introduces a set of low-level visual tools to integrate essential visual information into the reasoning chain, mitigating the neglect of fine-grained visual features. Second, a mask-based visual feedback mechanism is elaborated to incorporate visual reflection into the thinking process, enabling the model to dynamically re-examine and update its answers. Driven by RL, ForeSight learns to autonomously decide on tool invocation and answer verification, with the final answer accuracy as the reward signal. To evaluate the performance of the proposed framework, we construct a new dataset, Character and Grounding SalBench (CG-SalBench), based on the SalBench dataset. Experimental results demonstrate that the ForeSight-7B model significantly outperforms other models with the same parameter scale, and even surpasses the current SOTA closed-source models on certain metrics.

视觉推理强化学习多模态模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。