让智能体学会构建空间认知地图,提升长距离推理能力。
LASAR: Towards Spatio-temporal Reasoning with Latent Cognitive Map

- 设计双记忆架构,融合短期经验与长期语义地图
- 通过时空对比学习,实现2%-3.5%的零样本泛化提升
- 适合研究具身智能与空间推理的学者
具身人工智能的核心挑战在于验证智能体是否真正构建了空间结构的内部模型,而非仅模仿特定任务的专家轨迹。当前以动作为中心(如VLN)或推理为中心(如EQA)的方法普遍缺乏促使智能体编码长程、碎片化经验中精细空间关系(如拓扑或距离)的学习信号。为此,我们提出LASAR,一种具备双记忆系统的架构,用于同时维护情景记忆与语义认知地图。进一步引入时空上下文表示学习(ST-CRL),通过模拟中注释的时空上下文生成认知查询,构造样本对,从而从智能体经验中构建内部认知地图。实验表明,该方法在标准VLN-CE和VSI-Bench基准上实现2%-3.5%的零样本泛化性能提升,并证明所建认知地图具有高自一致性。
原文摘要 · Abstract (English)
A fundamental challenge in embodied AI is verifying if agents build internal models of spatial structure or merely learn to mimic task-specific expert trajectories. This is critical as foundational approaches rooted in action-centric tasks (e.g., VLN) and reasoning-centric tasks (e.g., EQA) often share a common limitation: they lack a learning signal that forces them to encode fine-grained spatial relationships (like topology or distance) over long-range, fragmented experiences. To address this, we first propose LASAR, an architecture featuring a dual-memory system designed to maintain both episodic experiences and a semantic cognitive map. We then introduce Spatio-temporal Contextual Representation Learning (ST-CRL), a contrastive objective designed to train this architecture. ST-CRL leverages spatio-temporal cues from cognitive queries generated through annotated spatio-temporal context in simulation to build sample pairs, thereby forming the internal cognitive map from the agent's experiences. Experiments demonstrate that our method achieves 2\%-3.5\% gains in both zero-shot generalization on standard VLN-CE and VSI-Bench benchmarks. We also demonstrate that our proposed cognitive map has high self-consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。