arXiv:2606.23634cs.CV2026-06中稿 · ECCV

无需3D模型,任意视角参考即可精准估计物体6D姿态

Pose Anything Anywhere:Model-free Object Poses from Arbitrary References

论文配图:Pose Anything Anywhere:Model-free Object Poses from Arbitrary References
图 1 · 摘自论文原文
  • 基于多视图变换器,学习跨视角几何一致性与对齐特征
  • 在YCB-V上比现有方法提升12%,LM-O上提升超20%
  • 支持单张或稀疏无姿态参考图,适合真实场景应用

估计未见过物体的6D姿态是开放世界机器人和具身感知中的基础挑战。基于模型的方法虽准确但依赖CAD资产或繁重的准备流程,而多数无模型方法仅限于成对单一锚点匹配,在遮挡和大视角变化下因查询-参考重叠度低而失效。为此,我们提出PANY,一种统一的无模型框架,可无缝处理RGB与RGB-D输入,支持单张或稀疏无姿态参考图,并有效泛化至新物体。基于多视图变压器几何主干,PANY突破成对匹配限制,学习在宽基线和有限重叠下仍稳定的视图一致几何与跨视图对齐线索。当存在额外无姿态辅助视图时,通过姿态图规范注册聚合它们,增强几何覆盖并强化最终姿态估计。大量实验表明,PANY在多个基准上达到当前最佳性能,显著优于现有无模型方法,在YCB-V上提升12%,在LM-O上提升超20%。且在单参考与稀疏参考设置下均表现稳定,展现强鲁棒性。

原文摘要 · Abstract (English)

Estimating the 6D pose of unseen objects is a fundamental yet challenging problem for open-world robotics and embodied perception. Model-based methods are accurate but depend on CAD assets or heavy onboarding, while most model-free approaches are still limited to pairwise single-anchor matching and thus fail under occlusion and large viewpoint changes with low query-reference overlap. Therefore, we present PANY, a unified model-free framework that seamlessly supports both RGB and RGB-D inputs, operates on one or sparse pose-free reference views, and generalizes effectively to novel objects. Built on a multi-view transformer geometry backbone, PANY moves beyond pairwise matching by learning view-consistent geometry and cross-view alignment cues that remain stable under wide baselines and limited overlap. When additional unposed assist views are available, PANY aggregates them via pose-graph canonical registration to increase geometric coverage and reinforce the final pose. Extensive experiments show that PANY achieves state-of-the-art performance across multiple benchmarks, substantially outperforming existing model-free methods, improving pose accuracy by +12% on YCB-V and over +20% on LM-O. Furthermore, PANY consistently performs well under both single-reference and sparse-reference settings, demonstrating strong robustness in real-world environments.

6D姿态估计无模型方法多视图融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。