统一人体姿态表示,用奇异值对比学习提升多模态对齐效果
UniHPR: Unified Human Pose Representation via Singular Value Contrastive Learning
- 提出基于奇异值的对比学习损失,同步对齐图像、2D/3D姿态等多模态表示
- 在Human3.6M上达49.9mm MPJPE,3DPW上51.6mm PA-MPJPE,跨域性能突出
- 支持2D/3D姿态检索,误差仅9.24mm MPJPE,适合多任务人体建模场景
近年来,构建跨模态统一表示以实现多模态融合与生成成为研究热点。作为以人为中心应用的关键组件,人体姿态表示在姿态估计、动作识别、人机交互、目标追踪等下游任务中至关重要。当前,从图像、2D关键点、3D骨架、网格模型等多种模态中提取的人体姿态嵌入尚未被系统性地通过对比学习方法进行关联研究。本文提出UniHPR,一种统一的人体姿态表示学习框架,可对齐图像、2D与3D人体姿态嵌入。为实现多于两种表示的同时对齐,我们设计了一种基于奇异值的对比学习损失,有效提升模态间对齐精度并进一步增强性能。我们在2D和3D人体姿态估计任务上评估所学表示:仅使用简单3D姿态解码器,UniHPR在Human3.6M数据集上取得49.9mm MPJPE,在3DPW数据集上达到51.6mm PA-MPJPE(跨域测试)。同时,利用统一表示在Human3.6M上实现了2D/3D姿态检索,平均误差为9.24mm MPJPE。
原文摘要 · Abstract (English)
In recent years, there has been a growing interest in developing effective alignment pipelines to generate unified representations from different modalities for multi-modal fusion and generation. As an important component of Human-Centric applications, Human Pose representations are critical in many downstream tasks, such as Human Pose Estimation, Action Recognition, Human-Computer Interaction, Object tracking, etc. Human Pose representations or embeddings can be extracted from images, 2D keypoints, 3D skeletons, mesh models, and lots of other modalities. Yet, there are limited instances where the correlation among all of those representations has been clearly researched using a contrastive paradigm. In this paper, we propose UniHPR, a unified Human Pose Representation learning pipeline, which aligns Human Pose embeddings from images, 2D and 3D human poses. To align more than two data representations at the same time, we propose a novel singular value-based contrastive learning loss, which better aligns different modalities and further boosts performance. To evaluate the effectiveness of the aligned representation, we choose 2D and 3D Human Pose Estimation (HPE) as our evaluation tasks. In our evaluation, with a simple 3D human pose decoder, UniHPR achieves remarkable performance metrics: MPJPE 49.9mm on the Human3.6M dataset and PA-MPJPE 51.6mm on the 3DPW dataset with cross-domain evaluation. Meanwhile, we are able to achieve 2D and 3D pose retrieval with our unified human pose representations in Human3.6M dataset, where the retrieval error is 9.24mm in MPJPE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。