让视觉语言模型学会动态更新信念,更像人一样推理。
Belief-Aware VLM Model for Human-like Reasoning

- 用向量记忆检索多模态上下文模拟人类信念
- 在HD-EPIC数据集上显著优于零样本基线
- 适合长期任务中的意图理解与决策场景
传统神经网络在意图推断中严重依赖可观测状态,难以跨任务和动态环境泛化。近期的视觉语言模型(VLM)和视觉语言动作(VLA)模型通过大规模多模态预训练引入常识推理,实现零样本跨任务性能。但这些模型缺乏显式表示和更新信念的机制,限制了其类人推理能力及对长期任务中意图演变的捕捉。为此,我们提出一种信念感知的VLM框架,融合基于检索的记忆系统与强化学习。不显式建模信念,而是用向量记忆检索相关多模态上下文,并融入VLM进行推理。进一步在VLM隐空间上使用强化学习策略优化决策。我们在公开的VQA数据集HD-EPIC上评估该方法,结果表明相比零样本基线有持续提升,凸显了信念感知推理的重要性。
原文摘要 · Abstract (English)
Traditional neural network models for intent inference rely heavily on observable states and struggle to generalize across diverse tasks and dynamic environments. Recent advances in Vision Language Models (VLMs) and Vision Language Action (VLA) models introduce common-sense reasoning through large-scale multimodal pretraining, enabling zero-shot performance across tasks. However, these models still lack explicit mechanisms to represent and update belief, limiting their ability to reason like humans or capture the evolving human intent over long-horizon. To address this, we propose a belief-aware VLM framework that integrates retrieval-based memory and reinforcement learning. Instead of learning an explicit belief model, we approximate belief using a vector-based memory that retrieves relevant multimodal context, which is incorporated into the VLM for reasoning. We further refine decision-making using a reinforcement learning policy over the VLM latent space. We evaluate our approach on publicly available VQA datasets such as HD-EPIC and demonstrate consistent improvements over zero-shot baselines, highlighting the importance of belief-aware reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。