用多尺度代码本协同优化动作与外观,生成更逼真的说话头视频。
Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video Generation
- 联合学习动作与外观代码本,实现跨尺度补偿。
- 在多个基准上优于现有方法,显著提升面部细节与姿态准确性。
- 适合需要高保真说话头生成的视频合成研究者。
说话头视频生成旨在从源图像保留人物身份,并从驱动视频中提取运动信息生成逼真视频。尽管该领域进展迅速,但同时精确建模复杂面部动作与细粒度外观仍具挑战。由于单张源图像难以提供充分外观指导,尤其在动态姿态变化下。为此,本文提出联合学习动作与外观代码本,并设计多尺度补偿机制,以有效优化解码过程中的运动条件与外观特征。具体地,在统一框架中同步学习多尺度动作与外观代码本,存储代表性全局面部运动流和外观模式;进一步提出基于Transformer的代码本检索策略,从两组代码本中查询互补信息,实现联合补偿。该流程生成更具灵活性的动作流与更少失真的外观特征,从而构建高质量说话头视频生成框架。大量实验验证了方法的有效性,在多个基准上均优于当前最优模型,从定性和定量角度均展现出优越性能。
原文摘要 · Abstract (English)
Talking head video generation aims to generate a realistic talking head video that preserves the person's identity from a source image and the motion from a driving video. Despite the promising progress made in the field, it remains a challenging and critical problem to generate videos with accurate poses and fine-grained facial details simultaneously. Essentially, facial motion is often highly complex to model precisely, and the one-shot source face image cannot provide sufficient appearance guidance during generation due to dynamic pose changes. To tackle the problem, we propose to jointly learn motion and appearance codebooks and perform multi-scale codebook compensation to effectively refine both the facial motion conditions and appearance features for talking face image decoding. Specifically, the designed multi-scale motion and appearance codebooks are learned simultaneously in a unified framework to store representative global facial motion flow and appearance patterns. Then, we present a novel multi-scale motion and appearance compensation module, which utilizes a transformer-based codebook retrieval strategy to query complementary information from the two codebooks for joint motion and appearance compensation. The entire process produces motion flows of greater flexibility and appearance features with fewer distortions across different scales, resulting in a high-quality talking head video generation framework. Extensive experiments on various benchmarks validate the effectiveness of our approach and demonstrate superior generation results from both qualitative and quantitative perspectives when compared to state-of-the-art competitors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。