arXiv:2512.15940cs.CVcs.RO2025-12被引 1

让视觉语言模型学会在四维时空里长期记忆与推理。

R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal Space

  • 用时空锚点构建可共享的持续世界模型。
  • 自然语言查询分解为语义、空间、时间三要素检索。
  • 无需训练即可实现情景化与协作式推理,适合机器人导航。

人类通过构建包含语义、空间布局和时间动态的持久结构化内部表征,在四维空间中感知与推理环境。受此启发,我们提出R4——一种无需训练的4D时空检索增强推理框架,使视觉语言模型(VLMs)具备结构化终身记忆。R4通过将物体级语义描述锚定于度量空间与时间,持续构建4D知识库,形成可跨智能体共享的持久世界模型。推理时,自然语言查询被分解为语义、空间和时间关键词以检索相关观测,并融入VLM推理过程。不同于传统检索增强生成方法,R4直接在4D空间中进行检索,实现无需训练的情景化与协作推理。在具身问答与导航基准测试中,R4显著优于基线,推动了动态环境中具身4D推理的新范式。

原文摘要 · Abstract (English)

Humans perceive and reason about their surroundings in four dimensions by building persistent, structured internal representations that encode semantic meaning, spatial layout, and temporal dynamics. These multimodal memories enable them to recall past events, infer unobserved states, and integrate new information into context-dependent reasoning. Inspired by this capability, we introduce R4, a training-free framework for retrieval-augmented reasoning in 4D spatio-temporal space that equips vision-language models (VLMs) with structured, lifelong memory. R4 continuously constructs a 4D knowledge database by anchoring object-level semantic descriptions in metric space and time, yielding a persistent world model that can be shared across agents. At inference, natural language queries are decomposed into semantic, spatial, and temporal keys to retrieve relevant observations, which are integrated into the VLM's reasoning. Unlike classical retrieval-augmented generation methods, retrieval in R4 operates directly in 4D space, enabling episodic and collaborative reasoning without training. Experiments on embodied question answering and navigation benchmarks demonstrate that R4 substantially improves retrieval and reasoning over spatio-temporal information compared to baselines, advancing a new paradigm for embodied 4D reasoning in dynamic environments.

视觉语言模型4D推理持续学习具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。