arXiv:2602.11073cs.CVcs.AI2026-02被引 1

让视觉模型通过对话动态调整图像理解,提升跨模态推理能力

Chatting with Images for Introspective Visual Thinking

论文配图:Chatting with Images for Introspective Visual Thinking
图 1 · 摘自论文原文
  • 用语言指令引导图像特征动态重编码,实现交互式视觉思考
  • 在8个基准上显著提升,尤其在多图与视频空间推理任务中表现突出
  • 适合需要复杂视觉逻辑推理的研究者和开发者使用

当前大型视觉语言模型(LVLMs)通常依赖单次视觉编码的纯文本推理,常导致细粒度视觉信息丢失。近期提出的“用图像思考”通过外部工具或代码操控图像缓解此问题,但生成的视觉状态往往缺乏语义支撑,难以实现有效的跨模态对齐——尤其在远距离区域或多图间的语义或几何关系推理时更为明显。为此,我们提出“与图像对话”框架,将视觉操作重构为语言引导的特征调制。在表达性语言提示指导下,模型动态对多个图像区域进行联合重编码,强化语言推理与视觉状态更新之间的耦合。我们在ViLaVT中实现了这一范式,这是一种专为交互式视觉推理设计的动态视觉编码器,并采用两阶段课程训练(监督微调+强化学习),以促进有效推理行为。在8个基准上的大量实验表明,ViLaVT取得显著且一致的性能提升,尤其在复杂的多图和视频空间推理任务中优势明显。

原文摘要 · Abstract (English)

Current large vision-language models (LVLMs) typically rely on text-only reasoning based on a single-pass visual encoding, which often leads to loss of fine-grained visual information. Recently the proposal of ''thinking with images'' attempts to alleviate this limitation by manipulating images via external tools or code; however, the resulting visual states are often insufficiently grounded in linguistic semantics, impairing effective cross-modal alignment - particularly when visual semantics or geometric relationships must be reasoned over across distant regions or multiple images. To address these challenges, we propose ''chatting with images'', a new framework that reframes visual manipulation as language-guided feature modulation. Under the guidance of expressive language prompts, the model dynamically performs joint re-encoding over multiple image regions, enabling tighter coupling between linguistic reasoning and visual state updates. We instantiate this paradigm in ViLaVT, a novel LVLM equipped with a dynamic vision encoder explicitly designed for such interactive visual reasoning, and trained it with a two-stage curriculum combining supervised fine-tuning and reinforcement learning to promote effective reasoning behaviors. Extensive experiments across eight benchmarks demonstrate that ViLaVT achieves strong and consistent improvements, with particularly pronounced gains on complex multi-image and video-based spatial reasoning tasks.

视觉语言模型交互推理多图理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。