TBDub提升视频配音的稳定性与效率,适合直播和生成视频场景。
TBDub: Production-Oriented Visual Dubbing
- 用生产数据微调+任务感知蒸馏,构建高效音画同步模型
- 学生模型实现7.13帧/秒,延迟降低13.93倍,保持高质量输出
- 适合需要实时性与高保真度的视频配音场景
视觉配音需同步口型与替换语音,同时保持身份、外观和时间一致性。尽管X-Dub提供了强大的无掩码视频编辑基线,但在直播和生成视频内容中仍存在生产域鲁棒性差、时间与运动不稳、身份和口腔细节保留不足及推理效率低的问题。我们提出TBDub,作为X-Dub的生产导向扩展,结合任务自适应后训练与任务感知少步蒸馏。后训练使用生产域数据、特定条件与过滤以及增强音频特征,得到30步教师模型。蒸馏将DMD/DMD2适配为条件视频编辑,并将教师压缩为两步学生模型。在38段TalkVid片段上,教师模型在所有八项重建、感知、身份与同步指标上优于X-Dub。MOS评估显示,其在唇同步一致性、身份一致性与视觉质量上分别提升0.14、0.95、0.90分;学生模型达到最高唇同步与视觉质量得分,身份一致性接近教师。在单块NVIDIA H20 GPU上,从首个VAE编码到最终解码的端到端生成耗时中,学生模型达7.13有效帧率,总延迟降低13.93倍;DiT阶段加速42.49倍。学生模型基本保留教师的生成质量与音画同步性能。代码已开源至GitHub,30步教师与两步学生权重可在Hugging Face获取。
原文摘要 · Abstract (English)
Visual dubbing must synchronize mouth motion with replacement speech while preserving identity, appearance, and temporal consistency. Although X-Dub provides a strong mask-free video-editing baseline, its application to livestream and generated-video content reveals limitations in production-domain robustness, temporal and motion stability, identity and oral-detail preservation, and inference efficiency. We present \textbf{TBDub}, a production-oriented extension of X-Dub that combines task-adaptive post-training with task-aware few-step distillation. Post-training adapts the video DiT using production-domain data, production-specific conditioning and filtering, and enhanced audio features to obtain a 30-step Teacher. Distillation adapts DMD/DMD2 to conditional video editing and compresses the Teacher into a two-step Student. On 38 TalkVid clips, the Teacher improves all eight reported reconstruction, perceptual, identity, and synchronization metrics over X-Dub. In the MOS evaluation, it improves lip-sync consistency, identity consistency, and visual quality over X-Dub by 0.14, 0.95, and 0.90 points, while the Student achieves the highest lip-sync and visual-quality scores and remains close to the Teacher in identity consistency. In paired end-to-end generation timing from the first VAE encode through the final VAE decode on a single NVIDIA H20 GPU at $512\times512$, the Student reaches 7.13 effective FPS and reduces total latency by $13.93\times$; the DiT stage alone is accelerated by $42.49\times$. The Student largely retains the Teacher's generation quality and audiovisual synchronization. The code is available on GitHub at \https://github.com/TaoLiveAIGC/TBDub, and the 30-step Teacher and two-step Student weights are available on Hugging Face at https://huggingface.co/TaoLiveAIGC/TBDub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。