arXiv:2604.12292cs.SDcs.CV2026-04被引 1

用认知同步扩散模型实现电影配音唇形精准对齐与自然音色还原

CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing

  • 基于认知模拟的扩散架构,分步优化音色、视觉对齐和时间上下文
  • 在标准与真实场景数据集上均达到当前最优性能,唇形同步率显著提升
  • 适合影视配音、虚拟人语音合成等需高保真同步的应用场景

电影配音旨在生成保留参考音频声线特征的同时与目标视频口型精确同步的语音。现有方法在时长层级显式对齐,难以实现精准唇形同步且自然度不足;虽有隐式对齐方案出现,但在真实场景下仍易受参考音频干扰,导致音色与发音质量下降。本文提出一种基于流匹配的电影配音框架——认知同步扩散变压器(CoSync-DiT),灵感来自专业演员的认知过程。该架构通过执行声学风格适配、细粒度视觉校准和时间感知上下文对齐,逐步引导从噪声到语音的生成轨迹。此外,设计联合语义与对齐正则化机制(JSAR),同时约束上下文输出的帧级时间一致性与流隐藏状态的语义一致性,确保鲁棒对齐。在标准基准与挑战性真实场景基准上的大量实验表明,本方法在多个指标上均达到当前最佳性能。

原文摘要 · Abstract (English)

Movie dubbing aims to synthesize speech that preserves the vocal identity of a reference audio while synchronizing with the lip movements in a target video. Existing methods fail to achieve precise lip-sync and lack naturalness due to explicit alignment at the duration level. While implicit alignment solutions have emerged, they remain susceptible to interference from the reference audio, triggering timbre and pronunciation degradation in in-the-wild scenarios. In this paper, we propose a novel flow matching-based movie dubbing framework driven by the Cognitive Synchronous Diffusion Transformer (CoSync-DiT), inspired by the cognitive process of professional actors. This architecture progressively guides the noise-to-speech generative trajectory by executing acoustic style adapting, fine-grained visual calibrating, and time-aware context aligning. Furthermore, we design the Joint Semantic and Alignment Regularization (JSAR) mechanism to simultaneously constrain frame-level temporal consistency on the contextual outputs and semantic consistency on the flow hidden states, ensuring robust alignment. Extensive experiments on both standard benchmarks and challenging in-the-wild dubbing benchmarks demonstrate that our method achieves the state-of-the-art performance across multiple metrics.

语音合成唇形同步扩散模型影视配音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。