arXiv:2609.08949cs.CV2026-09

无需已知物体模型,直接估计图像中多个未知物体的相对6D位姿。

Prior-free relative 6D pose estimation of multiple object instances

  • 利用多模态基础特征找物体实例间的粗略对应关系,无需训练。
  • 通过实例三元组的循环一致性约束优化对应关系,实现全局一致。
  • 在新构建的基准上超越现有方法,适合无先验场景下的多物体定位任务。

6D位姿估计正逐步减少对特定物体先验的依赖,从显式3D模型演进到多视角物体捕捉,再到单参考图像。本文将这一趋势推向极致,提出无先验相对6D位姿估计:不需知晓场景中具体是哪个物体,即可估计同一图像中多个未知物体的相对位姿。为此,我们设计了无需训练的PROSE方法,通过多模态基础特征建立物体实例间的粗略对应,再利用三元组循环一致性进行优化,最终基于全局一致的对应关系计算任意两实例间的相对6D位姿。为系统评估,我们构建了基于三个多实例BOP数据集的新基准PRENCH,包含任务特异性元数据。PROSE在该设置下持续优于适配的先进单图方法,且无需任务专属监督或额外学习模块。

原文摘要 · Abstract (English)

Object 6D pose estimation formulations have progressively reduced reliance on object-specific priors, evolving from explicit 3D models to multi-view object captures to single reference images. We take this progression to its extreme by introducing prior-free relative 6D pose estimation, which lifts the assumption of knowing which object is to be posed within the scene. This novel setting aims to estimate the relative poses of multiple instances of an unknown object within the same image, without requiring CAD models, templates, or reference images. We solve this by formulating a novel method (PROSE) that finds coarse correspondences between object instances using multimodal foundation features, thus requiring no training. We refine these correspondences by imposing cycle consistency across tuples of instances, and leverage the resulting globally consistent correspondences to estimate the relative 6D pose between any pair of instances. To enable systematic evaluation, we design a novel benchmark (PRENCH) built from three multi-instance BOP datasets and enriched with task-specific metadata. PROSE consistently outperforms baselines obtained by adapting state-of-the-art single-image methods to the proposed setting, while requiring neither task-specific supervision nor additional learned components. Project website: https://tev-fbk.github.io/PROSE/

6D位姿无先验多实例相对定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。