arXiv:2606.22076cs.CV2026-06

用跨视图语义先验提升单参考新物体位姿估计精度

Learning Cross-View Semantic Priors for Single-Reference Unseen Object Pose Estimation

论文配图:Learning Cross-View Semantic Priors for Single-Reference Unseen Object Pose Estimation
图 1 · 摘自论文原文
  • 引入跨视图语义交互,让查询与参考视图共享语义信息
  • 在六项基准上达到顶尖性能,尤其在挑战性视图对下表现优异
  • 适合需要快速部署新物体位姿估计的工业场景

单参考未见物体6D位姿估计算法通过仅需一个参考视图即可估计任意新物体的姿态,显著降低物体部署成本。现有基于对应关系的方法虽利用视觉基础模型(VFM)特征取得良好效果,但通常将这些特征视为视图内描述符,未能充分在视图间交换密集的视觉-语义线索(如外观、结构、上下文),导致解码后的点特征缺乏联合语义与几何区分性,在困难情况下仍难建立可靠对应。为此,本文构建早期跨视图语义先验,提出跨视图语义交互(CVSI),使查询与参考视图的VFM token可密集交换语义上下文并形成跨视图先验。为避免破坏原始特征结构并确保3D对应一致性,设计两个训练约束:视图内结构保持(IVSP)损失保留交互前的视图内亲和结构,参考锚定几何一致性(RAGC)损失强制解码点特征的空间表示一致性。最终通过加权SVD恢复姿态。我们在BOP挑战数据集YCB-V和TUD-L中构建了具有挑战性的视图对协议,实验表明,该方法在六个基准上均达最优性能,且推理速度相当。

原文摘要 · Abstract (English)

Single-reference unseen object 6D pose estimation reduces object onboarding by estimating poses of arbitrary novel objects from only one reference view. Recent correspondence-based pipelines have achieved robust performance with vision foundation model (VFM) features. However, they typically treat these features as intra-view descriptors, leaving dense visual-semantic cues, including appearance, structure, and context, insufficiently exchanged across views before geometric decoding. Consequently, the decoded point features may lack joint semantic and geometric discriminability, making correspondence estimation still difficult in challenging cases. Instead of processing features independently, we build the correspondence pipeline around an early cross-view semantic prior. Specifically, cross-view semantic interaction (CVSI) enables dense query and reference VFM tokens to exchange semantic context and form a cross-view prior. Nevertheless, direct CVSI may disturb the VFM token structure, while the resulting semantic prior still needs 3D representation consistency for rigid correspondence. To make this CVSI prior reliable for 3D correspondence learning, we introduce two complementary training-time constraints: the intra-view structure preservation (IVSP) loss preserves the original intra-view token affinity structure during interaction, while the reference-anchored geometric consistency (RAGC) loss enforces spatial representation consistency of decoded point features. The final pose is recovered from learned correspondences through weighted SVD. We further construct a challenging view-pair protocol from the BOP Challenge datasets YCB-V and TUD-L to evaluate robustness in difficult matching scenarios. Extensive experiments on six benchmarks under different view-pair settings show that our method achieves state-of-the-art performance while maintaining comparable inference speed.

位姿估计跨视图视觉基础模型6D姿态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。