用扩散模型提升单目相机对机器人位姿估计的鲁棒性
MonoSE(3)-Diffusion: A Monocular SE(3) Diffusion Framework for Robust Camera-to-Robot Pose Estimation
- 将位姿估计建模为带可见性约束的去噪扩散过程
- 在最难数据集上达66.75的AUC,比当前最优提升32.3%
- 适合需要高鲁棒性单目位姿估计的机器人应用
我们提出MonoSE(3)-Diffusion,一种基于单目图像的无标记机器人位姿估计扩散框架。该框架将位姿估计建模为条件去噪扩散过程,包含两个阶段:受可见性约束的扩散过程用于生成多样且在视场内的训练位姿,以及时间步感知的逆向过程实现逐步精炼。扩散过程通过逐步扰动真实位姿生成噪声变换以训练去噪网络,并引入可见性约束确保变换始终在相机视场内。相比现有方法固定的扰动尺度,本方法生成更具多样性且符合视场限制的训练样本,显著提升网络泛化能力。逆向过程通过去噪网络迭代预测位姿,并根据当前时间步的扩散后验采样进行优化,遵循粗到细的调度策略。时间步信息提示变换尺度,引导网络更精准预测。该方法在DREAM和RoboKeyGen两个基准上均取得提升,在最挑战的数据集上达到66.75的AUC,较当前最优提升32.3%。
原文摘要 · Abstract (English)
We propose MonoSE(3)-Diffusion, a monocular SE(3) diffusion framework that formulates markerless, image-based robot pose estimation as a conditional denoising diffusion process. The framework consists of two processes: a visibility-constrained diffusion process for diverse pose augmentation and a timestep-aware reverse process for progressive pose refinement. The diffusion process progressively perturbs ground-truth poses to noisy transformations for training a pose denoising network. Importantly, we integrate visibility constraints into the process, ensuring the transformations remain within the camera field of view. Compared to the fixed-scale perturbations used in current methods, the diffusion process generates in-view and diverse training poses, thereby improving the network generalization capability. Furthermore, the reverse process iteratively predicts the poses by the denoising network and refines pose estimates by sampling from the diffusion posterior of current timestep, following a scheduled coarse-to-fine procedure. Moreover, the timestep indicates the transformation scales, which guide the denoising network to achieve more accurate pose predictions. The reverse process demonstrates higher robustness than direct prediction, benefiting from its timestep-aware refinement scheme. Our approach demonstrates improvements across two benchmarks (DREAM and RoboKeyGen), achieving a notable AUC of 66.75 on the most challenging dataset, representing a 32.3% gain over the state-of-the-art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。