arXiv:2601.08175cs.CV2026-01被引 1

模仿人类认知,实现动态3D场景的持续记忆与快速重用

CogniMap3D: Cognitive 3D Mapping and Rapid Retrieval

  • 通过多阶段运动线索识别动态物体,结合深度和位姿先验
  • 建立持久化静态场景记忆库,支持跨次访问的快速检索与更新
  • 适用于长期连续环境感知,适合机器人导航与智能系统

我们提出CogniMap3D,一种受生物启发的动态3D场景理解与重建框架,模拟人类认知过程。该方法维护一个静态场景的持久记忆库,实现高效的空间知识存储与快速检索。CogniMap3D集成三项核心能力:基于多阶段运动线索的动态物体识别框架、支持跨多次访问的静态场景存储、回忆与更新的认知映射系统,以及用于优化相机位姿的因子图策略。给定图像流后,模型利用运动线索结合深度与相机位姿先验识别动态区域,再将静态元素与记忆库匹配。在重访熟悉位置时,CogniMap3D可检索存储场景,定位相机,并用新观测更新记忆。在视频深度估计、相机位姿重建与3D建图任务上的评估表明其达到当前最优性能,同时有效支持长时间序列与多次访问下的连续场景理解。

原文摘要 · Abstract (English)

We present CogniMap3D, a bioinspired framework for dynamic 3D scene understanding and reconstruction that emulates human cognitive processes. Our approach maintains a persistent memory bank of static scenes, enabling efficient spatial knowledge storage and rapid retrieval. CogniMap3D integrates three core capabilities: a multi-stage motion cue framework for identifying dynamic objects, a cognitive mapping system for storing, recalling, and updating static scenes across multiple visits, and a factor graph optimization strategy for refining camera poses. Given an image stream, our model identifies dynamic regions through motion cues with depth and camera pose priors, then matches static elements against its memory bank. When revisiting familiar locations, CogniMap3D retrieves stored scenes, relocates cameras, and updates memory with new observations. Evaluations on video depth estimation, camera pose reconstruction, and 3D mapping tasks demonstrate its state-of-the-art performance, while effectively supporting continuous scene understanding across extended sequences and multiple visits.

3D重建认知计算场景理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。