arXiv:2606.28026cs.CV2026-06中稿 · ECCV

解决人体动画中动作与形体纠缠问题,实现高保真表情与姿态控制。

EMOSH: Expressive Motion and Shape Disentanglement for Human Animation

论文配图:EMOSH: Expressive Motion and Shape Disentanglement for Human Animation
图 1 · 摘自论文原文
  • 提出显式分离形体与姿态参数的表达人体模型,根治形状泄露问题。
  • 在自驱动和跨驱动场景下均优于现有方法,生成视频细节丰富且身份一致。
  • 适合数字人、虚拟主播等需精细表情控制的应用场景。

高保真且富有表现力的可控人体动画对内容创作和数字人应用至关重要。然而,现有方法在表现力与解耦性之间面临两难:主流2D姿态条件方法存在“动作-形体纠缠”,导致驱动主体的体型信息泄露;依赖3D先验(如SMPL)的方法虽能实现几何解耦,却难以捕捉面部表情和复杂手势,造成动画僵硬。为此,我们提出EMOSH,一种用于高保真可控人体视频生成的新框架。首先,引入表达人体模型(EHM)作为核心控制表示,通过显式分离形体与姿态参数,从根本上解决体型泄露问题。同时设计鲁棒的动作追踪器,从视频中准确估计EHM参数。其次,提出粗到精的混合动作注入策略,实现对表情与手势的更细粒度控制。此外,引入空间对齐条件机制,弥合训练与推理间的域差距,提升身份一致性。大量实验表明,EMOSH在自驱动和跨驱动场景下均超越先前方法,生成具有生动表情的高质量视频,同时保持形体解耦。

原文摘要 · Abstract (English)

High-fidelity and expressive controllable human animation is essential for content creation and digital avatar applications. However, existing methods face a dilemma between expressiveness and disentanglement. Mainstream 2D pose-conditioned approaches suffer from "motion-shape entanglement", leading to the leakage of the driving subject's body shape. Conversely, methods relying on 3D priors (e.g., SMPL) achieve geometric disentanglement but struggle to capture facial expressions and complex gestures, resulting in rigid animations. To this end, we propose EMOSH, a novel framework for high-fidelity controllable human video generation. First, an Expressive Human Model (EHM) is introduced as the core control representation. By explicitly disentangling shape and pose parameters, we fundamentally resolve the body shape leakage issue. Alongside this, a robust motion tracker is designed to accurately estimate EHM parameters from video. Second, we propose a Coarse-to-Fine Hybrid Motion Injection strategy, enabling more fine-grained control over expressions and gestures. Furthermore, we introduce a Spatially-Aligned Conditioning mechanism to bridge the domain gap between training and inference, improving identity consistency. Extensive experiments demonstrate that EMOSH outperforms previous methods in both self-driven and cross-driven scenarios, producing high-fidelity videos with vivid expressions while maintaining shape disentanglement.

人体动画形体解耦表情控制视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。