arXiv:2603.14977cs.RO2026-03被引 1

用重投影点图融合视觉与几何信息,提升机器人操作的3D精度与泛化能力

ReMAP-DP: Reprojected Multi-view Aligned PointMaps for Diffusion Policy

  • 通过重投影点图实现多视角图像与3D结构对齐
  • 在RoboTwin 2.0上达59.3%成功率,比基线高6.6%
  • 仅需少量演示即可实现高精度真实世界操作

基于2D视觉表征的通用机器人策略在语义推理上表现优异,但缺乏高精度任务所需的显式3D空间感知。现有3D融合方法受限于稀疏点云的结构不规则性及多视图正交渲染带来的几何失真。为此,我们提出ReMAP-DP,结合标准化透视重投影与结构感知双流扩散策略。通过将重投影视图与像素对齐的PointMaps耦合,双流架构利用可学习模态嵌入融合冻结语义特征与显式几何描述符,实现精确的隐式块级对齐。大量仿真与真实环境实验表明,ReMAP-DP在多种操作任务中表现卓越:在RoboTwin 2.0上平均成功率达59.3%,较DP3基线提升6.6%;在ManiSkill 3的堆叠立方体任务中,几何挑战下性能相较DP3提升28%。此外,ReMAP-DP展现出显著的真实世界鲁棒性,仅需少量示范即可完成高精度动态操作,数据效率优异。

原文摘要 · Abstract (English)

Generalist robot policies built upon 2D visual representations excel at semantic reasoning but inherently lack the explicit 3D spatial awareness required for high-precision tasks. Existing 3D integration methods struggle to bridge this gap due to the structural irregularity of sparse point clouds and the geometric distortion introduced by multi-view orthographic rendering. To overcome these barriers, we present ReMAP-DP, a novel framework synergizing standardized perspective reprojection with a structure-aware dual-stream diffusion policy. By coupling the re-projected views with pixel-aligned PointMaps, our dual-stream architecture leverages learnable modality embeddings to fuse frozen semantic features and explicit geometric descriptors, ensuring precise implicit patch-level alignment. Extensive experiments across simulation and real-world environments demonstrate ReMAP-DP's superior performance in diverse manipulation tasks. On RoboTwin 2.0, it attains a 59.3% average success rate, outperforming the DP3 baseline by +6.6%. On ManiSkill 3, our method yields a 28% improvement over DP3 on the geometrically challenging Stack Cube task. Furthermore, ReMAP-DP exhibits remarkable real-world robustness, executing high-precision and dynamic manipulations with superior data efficiency from only a handful of demonstrations. Project page is available at: https://icr-lab.github.io/ReMAP-DP/

机器人操作扩散模型3D感知少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。