arXiv:2511.01501cs.CVcs.RO2025-11被引 7

用流模型估计6D位姿分布,让机器人更懂不确定性

SE(3)-PoseFlow: Estimating 6D Pose Distributions for Uncertainty-Aware Robotic Manipulation

  • 在SE(3)流形上做概率建模,生成多组可能位姿
  • 在Real275等数据集上精度领先,尤其在对称或遮挡场景
  • 适合需要判断置信度的机器人抓取与主动感知任务

物体位姿估计是机器人与计算机视觉中的基础问题,但受部分观测、遮挡和物体对称性影响,常导致位姿模糊与多解。现有确定性深度网络在约束良好时表现优异,却容易过度自信,无法捕捉位姿分布的多模态特性。为此,我们提出一种新的概率框架,在SE(3)流形上利用流匹配估计6D物体位姿分布。不同于仅输出单一确定结果的方法,本方法通过样本化方式建模完整位姿分布,可在对称物体或严重遮挡等模糊情形中进行不确定性推理。在Real275、YCB-V和LM-O数据集上达到当前最优性能,并验证了其在下游机器人操作任务中的应用价值,如通过主动感知消除视角不确定性,或以不确定性感知方式指导抓取生成。

原文摘要 · Abstract (English)

Object pose estimation is a fundamental problem in robotics and computer vision, yet it remains challenging due to partial observability, occlusions, and object symmetries, which inevitably lead to pose ambiguity and multiple hypotheses consistent with the same observation. While deterministic deep networks achieve impressive performance under well-constrained conditions, they are often overconfident and fail to capture the multi-modality of the underlying pose distribution. To address these challenges, we propose a novel probabilistic framework that leverages flow matching on the SE(3) manifold for estimating 6D object pose distributions. Unlike existing methods that regress a single deterministic output, our approach models the full pose distribution with a sample-based estimate and enables reasoning about uncertainty in ambiguous cases such as symmetric objects or severe occlusions. We achieve state-of-the-art results on Real275, YCB-V, and LM-O, and demonstrate how our sample-based pose estimates can be leveraged in downstream robotic manipulation tasks such as active perception for disambiguating uncertain viewpoints or guiding grasp synthesis in an uncertainty-aware manner.

6D位姿估计不确定性建模机器人抓取流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。