arXiv:2501.08408cs.CVcs.LG2025-01

用无标签图像提升3D姿态估计在跨域场景下的表现

Leveraging 2D Masked Reconstruction for Domain Adaptation of 3D Pose Estimation

  • 通过掩码图像建模,利用无标签数据增强模型泛化能力
  • 在跨域测试中显著提升准确率,达到当前最优水平
  • 适合需要应对真实世界复杂变化的3D姿态估计算法开发者

基于RGB的3D姿态估计方法在深度学习和高质量3D姿态数据集推动下取得了成功。然而,现有方法在测试图像分布与训练数据差异较大时表现不佳。虽然引入多样化数据可缓解此问题,但获取带标注(即3D姿态)的多样化数据极为困难。本文提出一种无监督域适应框架,通过掩码图像建模(MIM)框架融合无标签数据与有标签数据进行训练。进一步提出以前景为中心的重建和注意力正则化,提升无标签数据的利用效率。在人体与手部姿态估计任务的多个数据集上进行了实验,尤其针对跨域场景。结果表明,该方法在所有数据集上均达到了当前最优精度。

原文摘要 · Abstract (English)

RGB-based 3D pose estimation methods have been successful with the development of deep learning and the emergence of high-quality 3D pose datasets. However, most existing methods do not operate well for testing images whose distribution is far from that of training data. However, most existing methods do not operate well for testing images whose distribution is far from that of training data. This problem might be alleviated by involving diverse data during training, however it is non-trivial to collect such diverse data with corresponding labels (i.e. 3D pose). In this paper, we introduced an unsupervised domain adaptation framework for 3D pose estimation that utilizes the unlabeled data in addition to labeled data via masked image modeling (MIM) framework. Foreground-centric reconstruction and attention regularization are further proposed to increase the effectiveness of unlabeled data usage. Experiments are conducted on the various datasets in human and hand pose estimation tasks, especially using the cross-domain scenario. We demonstrated the effectiveness of ours by achieving the state-of-the-art accuracy on all datasets.

3D姿态估计域适应无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。