让机器人理解动态环境中的时空关系,提升导航与问答能力。
Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment
- 通过时空对齐构建以自我为中心的体验式世界模型。
- 在未知场景中实现1%-3%的准确率与探索效率提升。
- 适合研究具身智能、长程规划与视觉语言融合的学者。
在未知环境中实现类人推理仍是具身智能中的关键挑战。尽管先进视觉-语言模型(VLMs)在静态场景理解上表现优异,但在任务导向导航与具身问答(EQA)等动态、开放集任务中,仍受限于细粒度时空线索与物理世界理解建模不足。为此,我们提出VEME,一种新型跨模态对齐方法,通过学习以自我为中心、经验驱动的世界模型,增强未见场景的泛化能力。框架包含三部分:(1) 跨模态对齐机制,将物体、空间表征与视觉语义结合时空线索,提升VLM上下文学习能力;(2) 由世界嵌入激活的动态隐式认知地图,支持任务相关的几何-语义记忆召回;(3) 基于指令的导航与推理框架,利用具身先验实现长期规划与高效探索。通过嵌入几何感知的时空情景经验,该方法显著提升动态环境中的推理与规划性能。在VSI-Bench与VLN-CE数据集上的实验表明,相比传统方法,准确率与探索效率提升1%-3%。
原文摘要 · Abstract (English)
Achieving human-like reasoning in deep learning models for complex tasks in unknown environments remains a critical challenge in embodied intelligence. While advanced vision-language models (VLMs) excel in static scene understanding, their limitations in spatio-temporal reasoning and adaptation to dynamic, open-set tasks like task-oriented navigation and embodied question answering (EQA) persist due to inadequate modeling of fine-grained spatio-temporal cues and physical world comprehension. To address this, we propose VEME, a novel cross-modal alignment method that enhances generalization in unseen scenes by learning an ego-centric, experience-centered world model. Our framework integrates three key components: (1) a cross-modal alignment framework bridging objects, spatial representations, and visual semantics with spatio-temporal cues to enhance VLM in-context learning; (2) a dynamic, implicit cognitive map activated by world embedding to enable task-relevant geometric-semantic memory recall; and (3) an instruction-based navigation and reasoning framework leveraging embodied priors for long-term planning and efficient exploration. By embedding geometry-aware spatio-temporal episodic experiences, our method significantly improves reasoning and planning in dynamic environments. Experimental results on VSI-Bench and VLN-CE demonstrate 1%-3% accuracy and exploration efficiency improvement compared to traditional approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。