arXiv:2505.19049cs.CV2025-05

无监督学习下实现可控制的人体细粒度语义表示

Disentangled Human Body Representation Based on Unsupervised Semantic-Aware Learning

  • 设计骨架分组的全局感知解耦策略,建立几何语义与潜在编码对应关系
  • 在多个公开3D人体数据集上实现高精度重建,支持姿态迁移与双线性插值
  • 适合需要精细人体控制的应用,如动画生成、虚拟试衣

近年来,3D人体表征学习受到越来越多关注。然而,大量人工定义的人体约束和缺乏监督数据限制了现有方法在语义可控性和表征精度上的表现。本文提出一种基于无监督语义感知学习的人体表征方法,具备可控制的细粒度语义和高精度重建能力。特别地,设计了一种整体感知的骨架分组解耦策略,建立人体几何语义测量与潜在编码之间的对应关系,从而通过修改潜在参数实现对人体形状与姿态的精确控制。借助骨架分组的整体感知编码器和无监督解耦损失,模型在无监督条件下完成学习。此外,引入基于模板的残差学习方案,以缓解复杂身体形态与姿态空间中的潜在参数学习难度。由于潜在编码具有几何意义,该表示可广泛应用于人体姿态迁移、双线性潜在编码插值等任务。进一步采用部件感知解码器促进可控细粒度语义的学习。在多个公开3D人体数据集上的实验结果表明,该方法具备精确重建能力。

原文摘要 · Abstract (English)

In recent years, more and more attention has been paid to the learning of 3D human representation. However, the complexity of lots of hand-defined human body constraints and the absence of supervision data limit that the existing works controllably and accurately represent the human body in views of semantics and representation ability. In this paper, we propose a human body representation with controllable fine-grained semantics and high precison of reconstruction in an unsupervised learning framework. In particularly, we design a whole-aware skeleton-grouped disentangle strategy to learn a correspondence between geometric semantical measurement of body and latent codes, which facilitates the control of shape and posture of human body by modifying latent coding paramerers. With the help of skeleton-grouped whole-aware encoder and unsupervised disentanglement losses, our representation model is learned by an unsupervised manner. Besides, a based-template residual learning scheme is injected into the encoder to ease of learning human body latent parameter in complicated body shape and pose spaces. Because of the geometrically meaningful latent codes, it can be used in a wide range of applications, from human body pose transfer to bilinear latent code interpolation. Further more, a part-aware decoder is utlized to promote the learning of controllable fine-grained semantics. The experimental results on public 3D human datasets show that the method has the ability of precise reconstruction.

人体表征无监督学习解耦表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。