让视觉语言模型主动探索未知,用思维与视觉的差异驱动学习。
What You Think is What You See: Driving Exploration in VLM Agents via Visual-Linguistic Curiosity

- 用语言预测与视觉现实的差异作为探索驱动力。
- 在稀疏奖励任务中提升87%的完成率。
- 适合需要主动探索的复杂智能体任务。
为应对部分可观测的视觉环境,当前的视觉语言模型(VLM)智能体通过显式的思维链(CoT)推理将世界建模能力内化到策略中,能够在行动前进行未来模拟。然而,仅依赖已访问状态的被动推理不足以应对稀疏奖励任务,因缺乏主动发现‘未知之未知’的认知驱动力。本文提出GLANCE框架,通过将语言世界模型锚定在动态目标网络的稳定视觉表征上,实现推理与探索的统一。关键在于,利用语言预测与视觉现实之间的不一致作为强化学习中的内在好奇心信号,引导智能体主动探索内部模型不确定的区域。在一系列智能体任务上的实验表明,该方法显著有效,证明将‘智能体所想’与‘所见’对齐是解决复杂或稀疏任务的关键。
原文摘要 · Abstract (English)
To navigate partially observable visual environments, recent VLM agents increasingly internalize world modeling capabilities into their policies via explicit CoT reasoning, enabling them to mentally simulate futures before acting. However, relying solely on passive reasoning over visited states is insufficient for sparse-reward tasks, as it lacks the epistemic drive to actively uncover the ``known unknown'' required for robust generalization. We ask: Can VLM agents actively find signals that challenge and refine their internal world model through curiosity-driven exploration? In this work, we propose GLANCE, a unified framework that bridges reasoning and exploration by grounding the agent's linguistic world model into the stable visual representations of an evolving target network. Crucially, GLANCE leverages the discrepancy between linguistic prediction and visual reality as an intrinsic curiosity signal within reinforcement learning, steering the agent to actively explore areas where its internal model is uncertain. Extensive experiments across a series of agentic tasks show the effectiveness of GLANCE, and demonstrate that aligning ``what the agent thinks'' with ``what the agent sees'' is key to solving complex or sparse agentic tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。