用四组隐式编码实现跨身份高保真人体动画
X-UniMotion: Animating Human Images with Expressive, Unified and Identity-Agnostic Motion Latents

- 用四个解耦的隐式令牌分别表示面部、身体和双手动作
- 在多种身份和姿态下实现高保真动作迁移,无需显式骨骼参数
- 自监督训练+3D渲染增强,保持动作细节与身份一致性
我们提出X-UniMotion,一种统一且富有表现力的全身人体运动隐式表征,涵盖面部表情、身体姿态和手部动作。不同于依赖显式骨骼姿态和启发式跨身份调整的现有方法,本方法从单张图像直接编码多粒度运动为一组紧凑的四个解耦隐式令牌——一个用于面部表情,一个用于身体姿态,每个手一个。这些运动隐式令牌兼具高度表现力与身份无关性,可实现跨不同身份、姿态与空间配置的高保真、细节丰富的跨身份动作迁移。为此,我们设计了一个自监督端到端框架,联合学习运动编码器与隐式表征,并与基于DiT的视频生成模型协同训练,数据来自大规模多样的人体运动数据集。通过2D空间与色彩增强,以及共享姿态下的跨身份主体对合成3D渲染,强化运动-身份解耦。此外,借助辅助解码器引导运动令牌学习,促进细粒度、语义对齐且具有深度感知的嵌入。大量实验表明,X-UniMotion优于当前最优方法,在动作保真度与身份保留方面均表现更优。
原文摘要 · Abstract (English)
We present X-UniMotion, a unified and expressive implicit latent representation for whole-body human motion, encompassing facial expressions, body poses, and hand gestures. Unlike prior motion transfer methods that rely on explicit skeletal poses and heuristic cross-identity adjustments, our approach encodes multi-granular motion directly from a single image into a compact set of four disentangled latent tokens -- one for facial expression, one for body pose, and one for each hand. These motion latents are both highly expressive and identity-agnostic, enabling high-fidelity, detailed cross-identity motion transfer across subjects with diverse identities, poses, and spatial configurations. To achieve this, we introduce a self-supervised, end-to-end framework that jointly learns the motion encoder and latent representation alongside a DiT-based video generative model, trained on large-scale, diverse human motion datasets. Motion-identity disentanglement is enforced via 2D spatial and color augmentations, as well as synthetic 3D renderings of cross-identity subject pairs under shared poses. Furthermore, we guide motion token learning with auxiliary decoders that promote fine-grained, semantically aligned, and depth-aware motion embeddings. Extensive experiments show that X-UniMotion outperforms state-of-the-art methods, producing highly expressive animations with superior motion fidelity and identity preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。