首个基于离散流匹配的自动视频配音框架,实现精准口型同步与自然语音表达。
DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization
- 采用双阶段训练:先预训练零样本语音合成,再迁移至配音任务
- 通过跨模态对齐和面部表情引导,生成内容一致且口型同步的语音
- 适用于影视翻译、虚拟主播等需高保真语音同步的场景
视频配音需兼顾内容准确性、富有表现力的语调、高质量音效及精确的口型同步,现有方法在四方面均表现不佳。为此,我们提出DiFlowDubber,首个基于离散流匹配的视频配音框架,采用新颖的两阶段训练策略。第一阶段在大规模语料上预训练零样本文本到语音(TTS)系统,其中确定性架构捕捉语言结构,离散流式语调-声学(DFPA)模块建模富有表现力的语调与真实声学特征。第二阶段提出内容一致性时序适配(CCTA):其同步器强制实现跨模态对齐以实现口型同步;面部到语调映射器(FaPro)根据面部表情调节语调,输出与同步器融合,构建丰富细粒度的多模态嵌入,捕捉语调-内容关联,指导DFPA生成内容一致的语调与声学标记。在两个基准数据集上的实验表明,DiFlowDubber在多项评估指标上优于现有方法。
原文摘要 · Abstract (English)
Video dubbing requires content accuracy, expressive prosody, high-quality acoustics, and precise lip synchronization, yet existing approaches struggle on all four fronts. To address these issues, we propose DiFlowDubber, the first video dubbing framework built upon a discrete flow matching backbone with a novel two-stage training strategy. In the first stage, a zero-shot text-to-speech (TTS) system is pre-trained on large-scale corpora, where a deterministic architecture captures linguistic structures, and the Discrete Flow-based Prosody-Acoustic (DFPA) module models expressive prosody and realistic acoustic characteristics. In the second stage, we propose the Content-Consistent Temporal Adaptation (CCTA) to transfer TTS knowledge to the dubbing domain: its Synchronizer enforces cross-modal alignment for lip-synchronized speech. Complementarily, the Face-to-Prosody Mapper (FaPro) conditions prosody on facial expressions, whose outputs are then fused with those of the Synchronizer to construct rich, fine-grained multimodal embeddings that capture prosody-content correlations, guiding the DFPA to generate expressive prosody and acoustic tokens for content-consistent speech. Experiments on two benchmark datasets demonstrate that DiFlowDubber outperforms prior methods across multiple evaluation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。