arXiv:2603.00493cs.CV2026-03中稿 · ed被引 1

提出新方法,让单视图估计物体姿态更准更稳。

COG: Confidence-aware Optimal Geometric Correspondence for Unsupervised Single-reference Novel Object Pose Estimation

  • 用置信度引导的最优传输机制,生成平滑软对应关系。
  • 无监督下性能媲美有监督方法,有监督时更优。
  • 适合做无监督3D姿态估计的研究者和开发者。

仅凭单一参考视角估计新物体的6自由度位姿极具挑战,原因包括遮挡、视角变化和异常值。核心难点在于寻找鲁棒的跨视图对应关系,现有方法多依赖不可导的离散一对一匹配,易退化到稀疏关键点。本文提出置信度感知最优几何对应(COG),一种无监督框架,将对应关系估计建模为置信度感知的最优传输问题。COG通过预测点级置信度并将其作为最优传输边缘约束,生成平衡的软对应关系,抑制非重叠区域。视觉基础模型的语义先验进一步规整对应关系,实现稳定位姿估计。该设计将置信度融入对应与位姿估计全流程,支持无监督学习。实验表明,无监督COG性能可媲美有监督方法,有监督版本则超越它们。

原文摘要 · Abstract (English)

Estimating the 6DoF pose of a novel object with a single reference view is challenging due to occlusions, view-point changes, and outliers. A core difficulty lies in finding robust cross-view correspondences, as existing methods often rely on discrete one-to-one matching that is non-differentiable and tends to collapse onto sparse key-points. We propose Confidence-aware Optimal Geometric Correspondence (COG), an unsupervised framework that formulates correspondence estimation as a confidence-aware optimal transport problem. COG produces balanced soft correspondences by predicting point-wise confidences and injecting them as optimal transport marginals, suppressing non-overlapping regions. Semantic priors from vision foundation models further regularize the correspondences, leading to stable pose estimation. This design integrates confidence into the correspondence finding and pose estimation pipeline, enabling unsupervised learning. Experiments show unsupervised COG achieves comparable performance to supervised methods, and supervised COG outperforms them.

姿态估计无监督学习最优传输3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。