用3D视角对齐视觉与动作,提升机器人操作的泛化能力。
Dexterity-BEV: Aligning 3D World and Actions for Generalizable Robot Policies Learning

- 引入3D顶点图与谱表示,将2D图像升级为像素级3D信息。
- 构建统一鸟瞰视图框架,实现跨相机、机器人和数据集的对齐。
- 支持大规模训练与跨模态迁移,适合复杂场景下的机器人学习。
端到端操作策略结合网络规模预训练视觉语言模型,在通用且灵巧的机器人操作中展现出潜力。然而,它们继承了2D基础模型的两大局限:1)依赖2D RGB输入,忽略操作任务固有的3D特性;2)输入输出空间及不同机器人形态、相机设置、轨迹数据集间缺乏空间3D对齐。本文提出一系列贡献以解决上述问题。首先,引入对齐顶点图与顶点谱——一种基于相机标定与可选深度信息的像素级3D表示,将2D视觉输入提升至3D感知。其次,通过将每视角像素的3D信息与机器人动作映射到共享坐标系,实现输入输出对齐。在此基础上,定义规范鸟瞰视图(BEV)对齐帧,创新性地构建视图不变的BEV图像,增强对相机姿态变化的鲁棒性。为支持大规模训练与评估,开发完整的数据处理流水线,并引入新颖的时间对齐机制,实现跨机器人、人类操作者与数据集的轨迹对齐。这些贡献共同缓解了时空不对齐问题,显著提升真实场景下操作策略的一致性与泛化能力。预训练检查点、源代码与数据处理流水线已公开于 https://hnuzhy.github.io/projects/Dex-BEV。
原文摘要 · Abstract (English)
End-to-end manipulation policies, combined with web-scale pretrained Vision-Language Models (VLMs), show the promise for generalizable and dexterous robotic manipulation. However, they inherit two key limitations from 2D foundation models: 1) the reliance on 2D RGB inputs that ignores the intrinsically 3D nature of manipulation; and 2) the lack of spatial 3D alignment between input-output spaces as well as across diverse robot embodiments, camera setups, and trajectory datasets. In this paper, we present a series of contributions to address these issues. First, we introduce aligned vertex map and vertex spectrum -- a pixel-wise 3D representation that elevates 2D visual inputs to 3D, using camera calibration and optional depth. This novel input representation marries 3D awareness with the generalization of 2D large VLMs. Then, we propose to align the inputs and outputs of manipulation policies by expressing per-pixel 3D information of each camera view and robot actions to a shared coordinate. Based on this, we designate a canonical Bird's-Eye-View (BEV) alignment frame and innovatively propose to construct BEV images, producing a view-invariant representation robust to camera pose variations. To enable training and evaluation at scale, we develop a comprehensive data processing pipeline to perform such alignments; we also introduce a novel temporal alignment scheme for trajectories across diverse robots, human operators, and datasets. These contributions collectively mitigate input and output spatial-temporal misalignments, improving the consistency and generalization for real-world manipulation. Pretrained checkpoint, source code and data processing pipeline are available in https://hnuzhy.github.io/projects/Dex-BEV.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。