arXiv:2412.02066cs.CV2024-12

通过对比学习实现全范围头姿态估计,提升旋转外推能力。

CLERF: Contrastive LEaRning for Full Range Head Pose Estimation

  • 利用3D生成对抗网络生成三元组,解决数据稀疏难题。
  • 在轻微旋转/翻转图像上表现优于现有模型,全范围姿态预测准确。
  • 首个能准确估计倒置姿态的全范围头姿态模型,适合复杂场景应用。

我们提出一种新的头姿态估计(HPE)表征学习框架。以往因头姿态数据稀疏,难以进行三元组采样。近期3D生成对抗网络(3D-aware GAN)的发展使得三元组(锚点、正样本、负样本)采样变得容易。我们在大量增强数据(包括几何变换)上进行对比学习,证明该方法能使网络学习到真正有助于精确头姿态估计的特征。同时发现,现有HPE方法在测试图像旋转矩阵略微超出训练分布时性能下降明显。实验表明,我们的方法在标准测试集上达到顶尖水平,在轻微旋转或翻转图像及全范围头姿态下表现更优。据我们所知,这是首个能准确预测任意头姿态(包括倒置姿态)的真实全范围HPE模型。此外,与现有全偏航范围模型相比,结果更优。

原文摘要 · Abstract (English)

We introduce a novel framework for representation learning in head pose estimation (HPE). Previously such a scheme was difficult due to head pose data sparsity, making triplet sampling infeasible. Recent progress in 3D generative adversarial networks (3D-aware GAN) has opened the door for easily sampling triplets (anchor, positive, negative). We perform contrastive learning on extensively augmented data including geometric transformations and demonstrate that contrastive learning allows networks to learn genuine features that contribute to accurate HPE. On the other hand, we observe that existing HPE works struggle to predict head poses as accurately when test image rotation matrices are slightly out of the training dataset distribution. Experiments show that our methodology performs on par with state-of-the-art models on standard test datasets and outperforms them when images are slightly rotated/ flipped or full range head pose. To the best of our knowledge, we are the first to deliver a true full range HPE model capable of accurately predicting any head pose including upside-down pose. Furthermore, we compared with other existing full-yaw range models and demonstrated superior results.

头姿态估计对比学习3D生成全范围预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。