arXiv:2510.03885cs.RO2025-10被引 3

用3D隐空间地图提升机器人全局推理能力,让移动操作更智能。

Seeing the Bigger Picture: 3D Latent Mapping for Mobile Manipulation Policy Learning

  • 构建3D隐特征地图,融合多视角观测并扩展感知范围。
  • 在序列操作任务中成功率达90%,比图像基方法提升15%。
  • 适合需要长时记忆与全局规划的移动机器人任务。

本文展示,基于3D隐空间地图的移动操作策略在空间和时间推理能力上优于仅依赖图像的策略。我们提出端到端的SBP(Seeing the Bigger Picture)方法,直接在3D隐特征地图上运行。该地图通过增量融合多视角观测,构建场景特定的隐特征网格;使用预训练的场景无关解码器从这些特征重建目标嵌入,并支持任务执行中的在线优化。策略以隐地图为状态变量,通过3D特征聚合器获取全局上下文,可采用行为克隆或强化学习训练。我们在场景级移动操作和序列桌面操作任务上评估SBP,实验表明其(i)能全局推理场景,(ii)利用地图实现长时记忆,(iii)在分布内及新场景中均优于图像基策略,例如在序列操作任务中成功率提升15%。

原文摘要 · Abstract (English)

In this paper, we demonstrate that mobile manipulation policies utilizing a 3D latent map achieve stronger spatial and temporal reasoning than policies relying solely on images. We introduce Seeing the Bigger Picture (SBP), an end-to-end policy learning approach that operates directly on a 3D map of latent features. In SBP, the map extends perception beyond the robot's current field of view and aggregates observations over long horizons. Our mapping approach incrementally fuses multiview observations into a grid of scene-specific latent features. A pre-trained, scene-agnostic decoder reconstructs target embeddings from these features and enables online optimization of the map features during task execution. A policy, trainable with behavior cloning or reinforcement learning, treats the latent map as a state variable and uses global context from the map obtained via a 3D feature aggregator. We evaluate SBP on scene-level mobile manipulation and sequential tabletop manipulation tasks. Our experiments demonstrate that SBP (i) reasons globally over the scene, (ii) leverages the map as long-horizon memory, and (iii) outperforms image-based policies in both in-distribution and novel scenes, e.g., improving the success rate by 15% for the sequential manipulation task.

3D地图移动操作隐空间长时记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。