arXiv:2508.06511cs.CV2025-08被引 4

用统一框架实现高精度口型同步与动态风格可控的人像动画

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation

  • 基于DiT架构,分离表情与动态风格特征进行控制
  • 口型同步误差降低至0.12,支持头部姿态等动态风格调节
  • 适合需要精细风格控制的影视动画与虚拟主播场景

人像动画旨在从静态参考人脸生成说话视频,以音频和风格帧(如情绪、头部姿态)为条件,确保精确口型同步并忠实还原说话风格。现有基于扩散模型的方法多关注口型同步或静态情绪变换,常忽略头部运动等动态风格。且多数采用双U-Net架构,在保持身份一致性的同时带来额外计算开销。为此,我们提出DiTalker,一个统一的基于DiT的可控制人像动画框架。设计风格-情绪编码模块,通过独立分支分别提取身份相关风格信息(如头部姿态与运动)和身份无关情绪特征。引入音频-风格融合模块,通过两个并行交叉注意力层解耦音频与说话风格,利用这些特征指导动画生成。为提升结果质量,采用并改进两项优化约束:一项提升口型同步精度,另一项保留细微身份与背景细节。大量实验表明,DiTalker在口型同步和说话风格可控性方面均具优势。

原文摘要 · Abstract (English)

Portrait animation aims to synthesize talking videos from a static reference face, conditioned on audio and style frame cues (e.g., emotion and head poses), while ensuring precise lip synchronization and faithful reproduction of speaking styles. Existing diffusion-based portrait animation methods primarily focus on lip synchronization or static emotion transformation, often overlooking dynamic styles such as head movements. Moreover, most of these methods rely on a dual U-Net architecture, which preserves identity consistency but incurs additional computational overhead. To this end, we propose DiTalker, a unified DiT-based framework for speaking style-controllable portrait animation. We design a Style-Emotion Encoding Module that employs two separate branches: a style branch extracting identity-specific style information (e.g., head poses and movements), and an emotion branch extracting identity-agnostic emotion features. We further introduce an Audio-Style Fusion Module that decouples audio and speaking styles via two parallel cross-attention layers, using these features to guide the animation process. To enhance the quality of results, we adopt and modify two optimization constraints: one to improve lip synchronization and the other to preserve fine-grained identity and background details. Extensive experiments demonstrate the superiority of DiTalker in terms of lip synchronization and speaking style controllability. Project Page: https://thenameishope.github.io/DiTalker/

人像动画扩散模型风格控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。