无需预建地图,在动态室内环境实现可靠机器人操作与记忆更新。
Dynamic Resilient Spatio-Semantic Memory with Hybrid Localization for Mobile Manipulation

- 基于激光雷达-惯性-视觉融合的在线语义体素记忆构建
- 任务成功率提升至55%-70%,内存占用仅0.37-0.63GB
- 支持语言指令定位与物体重识别,适合真实场景机器人应用
在动态室内环境中实现可靠的移动操作,需要一种在环境变化时仍保持几何一致性、语义可查询且计算开销可控的场景表示。现有系统常依赖预建地图、静态场景假设或高精度相机位姿,导致目标物体移位或位姿校正后出现过时或错位的场景信息。本文提出DREAM框架,可在未见过的室内环境中实现感知、记忆、定位、导航与操作的端到端集成,无需预建地图。DREAM通过激光雷达-惯性-视觉SLAM后端注册的RGB-D观测构建在线语义体素记忆,并引入姿态图感知的冗余感知记忆剪枝(RMP)机制,在位姿修正后更新历史观测的同时,保持长时间观测历史的内存上限。针对目标定位与重捕获,DREAM融合语言条件3D检索、开放词汇图像检测及多模态大模型语义验证。四组动态实验室场景的真实机器人实验表明,相比DynaMem,DREAM将长时任务成功率从40%-60%提升至55%-70%,内存占用维持在0.37-0.63 GB,单次记忆更新耗时0.43-0.53秒。
原文摘要 · Abstract (English)
Reliable mobile manipulation in dynamic indoor environments requires a scene representation that remains geometrically consistent, semantically queryable, and computationally bounded as the environment changes. Existing systems often rely on pre-built maps, static-scene assumptions, or highly accurate camera poses, which can lead to stale or misaligned scene information when target objects are relocated or pose estimates are corrected. This paper presents DREAM, a real-robot mobile manipulation framework that integrates perception, memory, localization, navigation, and manipulation in previously unseen indoor environments without a pre-built map. DREAM constructs an online spatio-semantic voxel memory from RGB-D observations registered by a LiDAR-inertial-visual SLAM backend. It further introduces pose-graph-aware Redundancy-Aware Memory Pruning (RMP) to update historical observations after pose corrections while keeping long-horizon observation history bounded. For target localization and reacquisition, DREAM combines language-conditioned 3D retrieval, open-vocabulary image detection, and multimodal large language model based semantic verification. Real-robot experiments in four dynamic indoor laboratory scenes show that DREAM improves long-horizon task success rates from 40%-60% with DynaMem to 55%-70%, while maintaining a memory footprint of 0.37-0.63 GB and an online memory-update time of 0.43-0.53 s across scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。