arXiv:2510.18521cs.CV2025-10ICCV被引 6

用光线对齐思想提升未见物体6D姿态估计精度

RayPose: Ray Bundling Diffusion for Template Views in Unseen 6D Object Pose Estimation

  • 将姿态估计转为光线对齐问题,利用扩散Transformer建模
  • 在多个基准数据集上达到领先性能,尤其在未见物体上表现优异
  • 适合需要高精度6D姿态估计的机器人与AR应用

传统基于模板的物体姿态估计依赖匹配最相似模板并进行对齐,但模板检索错误常导致姿态预测不准。为此,本文将模板式姿态估计重构为光线对齐问题:学习多个带姿态模板图像的观察方向,使其与无姿态查询图像对齐。受基于扩散的相机姿态估计启发,采用扩散Transformer架构实现查询图与一组带姿态模板的对齐。通过物体中心相机光线重参数化旋转,并扩展尺度不变平移估计以建模密集平移偏移。模型利用模板中的几何先验引导查询姿态推理。基于缩小模板采样范围的粗到精训练策略提升了性能,且无需修改网络结构。在多个基准数据集上的大量实验表明,该方法在未见物体姿态估计任务中达到与最先进方法相当的效果。

原文摘要 · Abstract (English)

Typical template-based object pose pipelines estimate the pose by retrieving the closest matching template and aligning it with the observed image. However, failure to retrieve the correct template often leads to inaccurate pose predictions. To address this, we reformulate template-based object pose estimation as a ray alignment problem, where the viewing directions from multiple posed template images are learned to align with a non-posed query image. Inspired by recent progress in diffusion-based camera pose estimation, we embed this formulation into a diffusion transformer architecture that aligns a query image with a set of posed templates. We reparameterize object rotation using object-centered camera rays and model object translation by extending scale-invariant translation estimation to dense translation offsets. Our model leverages geometric priors from the templates to guide accurate query pose inference. A coarse-to-fine training strategy based on narrowed template sampling improves performance without modifying the network architecture. Extensive experiments across multiple benchmark datasets show competitive results of our method compared to state-of-the-art approaches in unseen object pose estimation.

6D姿态估计扩散模型视觉定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。