arXiv:2410.10122cs.CV2024-10被引 27

实时视频配音兼顾高画质与精准口型同步,突破现有技术瓶颈。

MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

论文配图:MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling
图 1 · 摘自论文原文
  • 分两阶段训练:先对齐面部姿态,再动态选择关键唇部区域优化。
  • 30帧/秒输出,256×256分辨率下在V100 GPU上实现实时推理。
  • 适合需要高质量实时配音的影视、直播及虚拟人应用。

实时视频配音在保持身份一致性的同时实现精准口型同步仍是重大挑战。现有方法面临三难困境:基于扩散模型的方法虽有高视觉保真度,但计算成本过高;而基于GAN的方案则牺牲口型同步精度或牙齿细节以达成实时性。本文提出MuseTalk,一种新型两阶段训练框架,通过潜空间优化与时空数据采样策略解决这一权衡。关键创新包括:(1) 在面部抽象预训练阶段,采用信息帧采样实现参考-源姿态的时间对齐,消除冗余特征干扰并保留身份线索;(2) 在唇形同步对抗微调阶段,使用动态边界采样法空间选取最有利于唇部运动的区域,平衡音视频同步与牙齿清晰度;(3) 在潜空间建立有效的音视频特征融合机制,实现在NVIDIA V100 GPU上以256×256分辨率生成30帧/秒的输出。大量实验表明,MuseTalk在视觉保真度上优于当前最优方法,同时保持相当的口型同步精度。

原文摘要 · Abstract (English)

Real-time video dubbing that preserves identity consistency while achieving accurate lip synchronization remains a critical challenge. Existing approaches face a trilemma: diffusion-based methods achieve high visual fidelity but suffer from prohibitive computational costs, while GAN-based solutions sacrifice lip-sync accuracy or dental details for real-time performance. We present MuseTalk, a novel two-stage training framework that resolves this trade-off through latent space optimization and spatio-temporal data sampling strategy. Our key innovations include: (1) During the Facial Abstract Pretraining stage, we propose Informative Frame Sampling to temporally align reference-source pose pairs, eliminating redundant feature interference while preserving identity cues. (2) In the Lip-Sync Adversarial Finetuning stage, we employ Dynamic Margin Sampling to spatially select the most suitable lip-movement-promoting regions, balancing audio-visual synchronization and dental clarity. (3) MuseTalk establishes an effective audio-visual feature fusion framework in the latent space, delivering 30 FPS output at 256*256 resolution on an NVIDIA V100 GPU. Extensive experiments demonstrate that MuseTalk outperforms state-of-the-art methods in visual fidelity while achieving comparable lip-sync accuracy. %The codes and models will be made publicly available upon acceptance. The code is made available at \href{https://github.com/TMElyralab/MuseTalk}{https://github.com/TMElyralab/MuseTalk}

视频配音口型同步实时生成潜空间优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。