无需人脸裁剪或说话人分离,实现多说话人对话的精准视频配音。
CineDub: Scaling End-to-End Video Dubbing to Multi-Speaker Dialogues with Coherent Sound Effects

- 采用统一扩散模型直接处理未裁剪视频,通过跨模态训练隐式关联视觉与语义信息。
- 在多说话人对话和音效同步生成上达到顶尖性能,首次实现语音与环境音联合生成。
- 适用于影视字幕本地化、多角色动画配音等真实场景,尤其适合复杂对话环境。
真实场景下的自动视频配音仍受限于两个相互冲突的约束:分层方法依赖脆弱的多阶段预处理流程,严重限制数据可扩展性和实际部署;而端到端方法在未裁剪视频上运行时,在多说话人场景中存在时间对齐弱和说话人-话语模糊的问题。为此,我们提出 CineDub,一种基于扩散模型的统一框架,可直接从未裁剪视频中实现精确的多说话人对话配音,无需人脸裁剪或说话人分离。核心是隐式耦合的整体条件(ICHC)范式,其中整体视觉表示与语义打包的转录文本独立编码,但通过跨模态训练隐式耦合,以解决说话人模糊性并实现精准的多说话人多轮对话配音。基于整体视觉特征捕捉的统一时间线索,我们进一步将 CineDub 扩展至语音与音频联合生成。引入环境到语言的课程学习(ALC)缓解子任务退化,并设计解耦文本分支控制机制以消除同步生成中的跨提示干扰。我们还发布了两个真实场景基准:CineDub-Multi 用于多说话人对话配音,CineDub-SA 用于视频到语音与音频(V2SA)生成。实验表明,CineDub 在既有单说话人配音与视频到音频基准上达到最先进水平,且在多说话人对话配音与声学一致的联合生成中表现卓越。
原文摘要 · Abstract (English)
Automatic video dubbing in the wild remains fundamentally limited by two competing constraints: hierarchical methods depend on brittle, multi-stage preprocessing pipelines that severely restrict data scalability and practical deployment, while holistic approaches operating on uncropped video suffer from weak temporal alignment and speaker-utterance ambiguity in multi-speaker settings. To overcome these limitations, we propose CineDub, a unified diffusion-based model that achieves precise multi-speaker dialogue dubbing directly from uncropped videos, without face cropping or speaker diarization. Central to our approach is the Implicitly-Coupled Holistic Conditioning (ICHC) paradigm, where holistic visual representations and a semantic-bundled transcription format are encoded independently, yet implicitly coupled through cross-modal training to resolve speaker ambiguity and enable precise multi-speaker multi-turn dialogue dubbing. Building on the unified temporal cues captured by holistic visual features, we further extend CineDub to joint speech and audio generation. We introduce an Ambient-to-Linguistic Curriculum Learning (ALC) to mitigate sub-task degradation, and a decoupled textual branch control mechanism to resolve cross-prompt interference during simultaneous generation. We also release two in-the-wild benchmarks, CineDub-Multi for multi-speaker dialogue dubbing and CineDub-SA for video-to-speech-and-audio (V2SA) generation, to enable evaluation under realistic conditions. Experiments show that CineDub achieves state-of-the-art results on established single-speaker dubbing and video-to-audio benchmarks while excelling in multi-speaker dialogue dubbing and acoustically coherent joint generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。