arXiv:2602.23937cs.ROcs.CV2026-02被引 2

用真实室内视频构建多模态事件知识图谱,提升导航模型长程推理能力。

Enhancing Vision-Language Navigation with Multimodal Event Knowledge from Real-World Indoor Tour Videos

  • 从真实视频中提取语义-动作-效果事件,构建首个大规模多模态时空知识图谱
  • 在三个基准上超越现有方法,尤其在复杂指令和长路径导航中表现显著
  • 适合研究视觉语言导航、具身智能与多模态知识融合的学者与工程师

视觉语言导航(VLN)代理在未见环境中的长程推理常因模糊、粗粒度指令而受阻。尽管已有研究利用知识图谱增强推理,但受人类情景记忆启发的多模态事件知识仍未被充分挖掘。本文提出一种以事件为中心的知识增强策略,用于自动化过程知识挖掘与特征融合,解决VLN任务中的粗粒度指令与长程推理难题。首先,我们基于真实室内视频构建了首个大规模多模态时空知识图谱YE-KG,包含超过86,000个节点和83,000条边;通过LLaVa、GPT4等多模态大模型,将非结构化视频流转化为结构化的语义-动作-效果事件,作为显式情景记忆。其次,我们提出STE-VLN,采用粗到细分层检索机制,将知识图谱中的因果事件序列与自中心视觉观测动态融合。在REVERIE、R2R和R2R-CE基准上的实验表明,该策略在多种动作空间下均优于当前最优方法。数据与代码已公开于项目网站 https://sites.google.com/view/y-event-kg/。

原文摘要 · Abstract (English)

Vision-Language Navigation (VLN) agents often struggle with long-horizon reasoning in unseen environments, particularly when facing ambiguous, coarse-grained instructions. While recent advances use knowledge graph to enhance reasoning, the potential of multimodal event knowledge inspired by human episodic memory remains underexplored. In this work, we propose an event-centric knowledge enhancement strategy for automated process knowledge mining and feature fusion to solve coarse-grained instruction and long-horizon reasoning in VLN task. First, we construct YE-KG, the first large-scale multimodal spatiotemporal knowledge graph, with over 86k nodes and 83k edges, derived from real-world indoor videos. By leveraging multimodal large language models (i.e., LLaVa, GPT4), we extract unstructured video streams into structured semantic-action-effect events to serve as explicit episodic memory. Second, we introduce STE-VLN, which integrates the above graph into VLN models via a Coarse-to-Fine Hierarchical Retrieval mechanism. This allows agents to retrieve causal event sequences and dynamically fuse them with egocentric visual observations. Experiments on REVERIE, R2R, and R2R-CE benchmarks demonstrate the efficiency of our event-centric strategy, outperforming state-of-the-art approaches across diverse action spaces. Our data and code are available on the project website https://sites.google.com/view/y-event-kg/.

视觉语言导航事件知识多模态知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。