通过局部特征对齐与伪标签优化,提升模拟到真实环境的物体姿态估计性能。
Unsupervised Domain Adaptation for Sim-to-Real Object Pose Estimation with Contrastive Alignment and Pseudo-Label Refinement

- 基于中间特征匹配无监督定位相似图像对,实现跨域配对。
- 在局部区域进行对比对齐,保留关键几何信息以提升姿态精度。
- 利用一致性约束优化伪标签,增强目标域预测稳定性。
无监督域适应(UDA)可将模拟环境中的知识鲁棒迁移至真实场景,仅使用少量未标注的真实数据提升性能。现有物体姿态估计的UDA方法多依赖全局特征匹配、多阶段大模型或图像翻译流程,常忽略特征表示中嵌入的姿态相关线索。为此,本文提出CAPLR,专注于局部区域中姿态敏感特征的适应,确保域对齐过程保留对准确姿态估计至关重要的几何信息。CAPLR包含三个核心组件:(1) 高效跨域配对策略,利用中间特征在无监督条件下识别跨域姿态相似的图像对;(2) 对比对齐机制,在中间特征和任务特定表示的局部区域实现特征对齐;(3) 基于一致性的伪标签精炼,通过鼓励目标域预测稳定来提升可靠性。大量实验表明,CAPLR在多个知名物体姿态估计基准上表现领先,涵盖多样且具有挑战性的场景。
原文摘要 · Abstract (English)
Unsupervised domain adaptation (UDA) enables robust transfer of knowledge from simulated to real environments while exploiting a subset of unlabeled target data to improve real-world performance. Existing UDA methods for Object pose estimation often rely on global feature matching, multi-stage larger frameworks, or image translation pipelines, which tend to overlook the pose-specific information embedded in feature representations. To bridge this limitation, we introduce CAPLR that targets the adaptation of pose-sensitive features in localized regions, ensuring that domain alignment preserves the geometric cues essential for accurate pose estimation. CAPLR achieves UDA with three key components: (1) Efficient Cross-Domain Pairing strategy leveraging intermediate features to identify pose similar image pairs across domains without supervision; (2) Contrastive Alignment to perform feature alignment at localised regions in both intermediate and task-specific representations; and (3) Consistency-Based Pseudo-Label Refinement to improve reliability by encouraging stable target predictions. Extensive experiments demonstrate that CAPLR achieves state-of-the-art performance across multiple well-known object pose estimation benchmarks featuring diverse and challenging scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。