arXiv:2602.02873cs.CV2026-02

让视觉语言模型像人一样主动提问,精准获取关键视觉信息。

ViThinker: Active Vision-Language Reasoning via Dynamic Perceptual Querying

  • 通过动态生成查询令牌,主动触发专家级视觉特征合成。
  • 在多个视觉基准上提升感知锚定与推理准确率,优于被动方法。
  • 适合需要精细视觉理解的复杂多模态任务研究者使用。

链式思维(CoT)推理在语言模型中表现优异,但在视觉语言模型中受限于过早的视觉转文本过程,导致几何与空间布局等连续信息丢失。现有方法虽通过静态枚举或基于注意力的选择改进CoT,但仍为被动处理预计算输入。受人类主动感知启发,我们提出ViThinker框架,使视觉语言模型能自主生成决策(查询)令牌,按需触发专家对齐的视觉特征合成。训练时内化视觉专家能力,推理时无需外部工具调用即可进行生成式心理模拟。采用两阶段课程:首先将冻结专家知识蒸馏至模型参数,再通过稀疏性惩罚学习任务驱动的查询策略,使ViThinker发现每个推理步骤所需的最小充分感知。在多个以视觉为中心的基准上评估显示一致提升,验证了主动查询生成在感知锚定与推理准确性上均优于被动方法。

原文摘要 · Abstract (English)

Chain-of-Thought (CoT) reasoning excels in language models but struggles in vision-language models due to premature visual-to-text conversion that discards continuous information such as geometry and spatial layout. While recent methods enhance CoT through static enumeration or attention-based selection, they remain passive, i.e., processing pre-computed inputs rather than actively seeking task-relevant details. Inspired by human active perception, we introduce ViThinker, a framework that enables vision-language models to autonomously generate decision (query) tokens triggering the synthesis of expert-aligned visual features on demand. ViThinker internalizes vision-expert capabilities during training, performing generative mental simulation during inference without external tool calls. Through a two-stage curriculum: first distilling frozen experts into model parameters, then learning task-driven querying via sparsity penalties, i.e., ViThinker discovers minimal sufficient perception for each reasoning step. Evaluations across vision-centric benchmarks demonstrate consistent improvements, validating that active query generation outperforms passive approaches in both perceptual grounding and reasoning accuracy.

视觉推理主动感知链式思维多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。