让大模型像人一样看图思考,实现深度交互式视觉推理。
V-Thinker: Interactive Thinking with Images
- 通过强化学习构建端到端的图像交互思考框架。
- 在多个任务上超越现有大模型基线,尤其在复杂推理中表现突出。
- 适合需要精细图像分析与长期推理的应用场景。
让大型多模态模型(LMMs)深度融合图像交互与长时程推理能力,仍是该领域长期面临的挑战。近期视觉中心的推理研究提出了‘看图思考’的新范式,推动模型从图像辅助推理转向图像交互式思考。尽管这一进展使模型能聚焦图像细粒度区域,但进展仍受限于有限的视觉工具空间和任务特定的工作流设计。为此,我们提出V-Thinker,一个通用型多模态推理助手,通过端到端强化学习实现交互式、以视觉为中心的思考。V-Thinker包含两个核心组件:(1) 数据演化飞轮,可自动合成、演化并验证跨多样性、质量、难度三个维度的交互式推理数据集;(2) 视觉渐进式训练课程,先通过点级监督对齐感知,再通过两阶段强化学习整合交互式推理。此外,我们引入VTBench——一个专家验证的基准,专门针对视觉中心的交互式推理任务。大量实验表明,V-Thinker在通用与交互式推理场景中均持续优于强基线的LMM,为推进图像交互式推理应用提供了重要启示。
原文摘要 · Abstract (English)
Empowering Large Multimodal Models (LMMs) to deeply integrate image interaction with long-horizon reasoning capabilities remains a long-standing challenge in this field. Recent advances in vision-centric reasoning explore a promising "Thinking with Images" paradigm for LMMs, marking a shift from image-assisted reasoning to image-interactive thinking. While this milestone enables models to focus on fine-grained image regions, progress remains constrained by limited visual tool spaces and task-specific workflow designs. To bridge this gap, we present V-Thinker, a general-purpose multimodal reasoning assistant that enables interactive, vision-centric thinking through end-to-end reinforcement learning. V-Thinker comprises two key components: (1) a Data Evolution Flywheel that automatically synthesizes, evolves, and verifies interactive reasoning datasets across three dimensions-diversity, quality, and difficulty; and (2) a Visual Progressive Training Curriculum that first aligns perception via point-level supervision, then integrates interactive reasoning through a two-stage reinforcement learning framework. Furthermore, we introduce VTBench, an expert-verified benchmark targeting vision-centric interactive reasoning tasks. Extensive experiments demonstrate that V-Thinker consistently outperforms strong LMM-based baselines in both general and interactive reasoning scenarios, providing valuable insights for advancing image-interactive reasoning applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。