让虚拟人面部与动作同步生成,提升数字人表现力。
AsynFusion: Towards Asynchronous Latent Consistency Models for Decoupled Whole-Body Audio-Driven Avatars
- 双分支扩散变换器并行生成表情与动作
- 协同同步模块实现跨模态特征交互
- 异步采样策略降低计算开销,实时性更强
全身音频驱动的虚拟人姿态与表情生成是打造逼真数字人和增强互动虚拟代理能力的关键任务,广泛应用于虚拟现实、数字娱乐和远程通信。现有方法通常独立生成面部表情与手势,导致二者缺乏自然协调,动画生硬不连贯。为此,我们提出 AsynFusion,一种基于双分支扩散变换器(DiT)的新框架,实现表情与动作的和谐合成。模型引入协同同步模块,促进两模态间双向特征交互;采用异步潜空间一致性模型(Asynchronous LCM)采样策略,在保持高质量输出的同时显著降低计算开销。大量实验表明,AsynFusion 在生成实时、同步的全身动画方面达到当前最优性能,在定量与定性评估中均优于现有方法。
原文摘要 · Abstract (English)
Whole-body audio-driven avatar pose and expression generation is a critical task for creating lifelike digital humans and enhancing the capabilities of interactive virtual agents, with wide-ranging applications in virtual reality, digital entertainment, and remote communication. Existing approaches often generate audio-driven facial expressions and gestures independently, which introduces a significant limitation: the lack of seamless coordination between facial and gestural elements, resulting in less natural and cohesive animations. To address this limitation, we propose AsynFusion, a novel framework that leverages diffusion transformers to achieve harmonious expression and gesture synthesis. The proposed method is built upon a dual-branch DiT architecture, which enables the parallel generation of facial expressions and gestures. Within the model, we introduce a Cooperative Synchronization Module to facilitate bidirectional feature interaction between the two modalities, and an Asynchronous LCM Sampling strategy to reduce computational overhead while maintaining high-quality outputs. Extensive experiments demonstrate that AsynFusion achieves state-of-the-art performance in generating real-time, synchronized whole-body animations, consistently outperforming existing methods in both quantitative and qualitative evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。