让虚拟人说话更像本人,还能精准对口型。
MIRRORTALK: Forging Personalized Avatars Via Disentangled Style and Hierarchical Motion Control
- 用解耦风格编码器从短视频提取说话风格特征
- 通过分层调制实现口型同步与面部表情的双重优化
- 适合需要高度个性化虚拟形象的视频生成场景
生成既保留说话者独特风格又精准对口型的个性化虚拟人脸仍是难题。现有方法常将说话风格与语义内容混淆,难以忠实迁移个人特征。本文提出MirrorTalk,基于条件扩散模型,结合语义解耦风格编码器(SDSE),可从简短视频中提取纯净风格表征。进一步在扩散过程中引入分层调制策略,动态平衡音频与风格特征在不同面部区域的贡献,确保口型精准且面部表情生动。大量实验表明,该方法在口型同步准确性和个性化保留方面显著优于现有最佳方法。
原文摘要 · Abstract (English)
Synthesizing personalized talking faces that uphold and highlight a speaker's unique style while maintaining lip-sync accuracy remains a significant challenge. A primary limitation of existing approaches is the intrinsic confounding of speaker-specific talking style and semantic content within facial motions, which prevents the faithful transfer of a speaker's unique persona to arbitrary speech. In this paper, we propose MirrorTalk, a generative framework based on a conditional diffusion model, combined with a Semantically-Disentangled Style Encoder (SDSE) that can distill pure style representations from a brief reference video. To effectively utilize this representation, we further introduce a hierarchical modulation strategy within the diffusion process. This mechanism guides the synthesis by dynamically balancing the contributions of audio and style features across distinct facial regions, ensuring both precise lip-sync accuracy and expressive full-face dynamics. Extensive experiments demonstrate that MirrorTalk achieves significant improvements over state-of-the-art methods in terms of lip-sync accuracy and personalization preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。