让视频配音时全身动作自然同步声音,支持无限长视频生成
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
- 用关键帧保留身份和动作轨迹,实现全身同步配音
- 在三个数据集上达到当前最优,动作与声音高度同步
- 适合需要长期口型同步的影视配音与虚拟人应用
近期视频AIGC突破推动了音频驱动的人类动画发展。但传统配音技术仅限于嘴部区域编辑,导致面部表情和身体动作不协调,影响观感沉浸。为此,我们提出稀疏帧视频配音新范式,通过有策略地保留参考关键帧,维持身份、标志性动作和镜头轨迹,同时实现全身动作的音画同步编辑。通过深入分析,我们发现朴素的图像到视频模型在此任务中失败,主要因其无法实现自适应条件控制。针对此问题,我们提出InfiniteTalk,一种面向无限长序列配音的流式音频驱动生成器。该架构利用时间上下文帧实现片段间无缝过渡,并引入简单有效的采样策略,通过精细定位参考帧优化控制强度。在HDTF、CelebV-HQ和EMTD数据集上的全面评估表明,其性能达当前最优水平。定量指标证实其在视觉真实感、情感一致性及全身动作同步性方面显著领先。
原文摘要 · Abstract (English)
Recent breakthroughs in video AIGC have ushered in a transformative era for audio-driven human animation. However, conventional video dubbing techniques remain constrained to mouth region editing, resulting in discordant facial expressions and body gestures that compromise viewer immersion. To overcome this limitation, we introduce sparse-frame video dubbing, a novel paradigm that strategically preserves reference keyframes to maintain identity, iconic gestures, and camera trajectories while enabling holistic, audio-synchronized full-body motion editing. Through critical analysis, we identify why naive image-to-video models fail in this task, particularly their inability to achieve adaptive conditioning. Addressing this, we propose InfiniteTalk, a streaming audio-driven generator designed for infinite-length long sequence dubbing. This architecture leverages temporal context frames for seamless inter-chunk transitions and incorporates a simple yet effective sampling strategy that optimizes control strength via fine-grained reference frame positioning. Comprehensive evaluations on HDTF, CelebV-HQ, and EMTD datasets demonstrate state-of-the-art performance. Quantitative metrics confirm superior visual realism, emotional coherence, and full-body motion synchronization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。