arXiv:2601.22143cs.GRcs.CV2026-01International Conf…被引 5

用联合音视频扩散模型实现一键多语言配音,保持口型同步与说话人特征。

JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion

  • 基于音视频扩散模型,通过轻量LoRA适配实现端到端配音。
  • 合成双语视频并互相修复音视频,训练出跨语言同步能力。
  • 适合需要高质量口型同步的影视本地化场景。

音频-视觉基础模型能够联合生成声音与视觉内容,展现出前所未有的多模态生成与编辑能力,为下游任务带来新机遇。在这些任务中,视频配音可显著受益于此类先验知识,但现有方案大多依赖复杂、特定任务的流水线,在真实场景中表现不佳。本文提出一种单模型方法,通过轻量级LoRA适配一个基础音视频扩散模型,实现从视频到视频的配音。该方法使模型能以输入音视频为条件,联合生成翻译后的音频与同步的面部动作。为训练此LoRA,我们利用生成模型自身合成同一说话人的多语言视频对:在单个视频片段内进行语言切换,然后分别对半段画面与音频进行修补,使其与另一半的语言一致。借助音视频模型丰富的生成先验,该方法在保留说话人身份的同时,维持了良好的口型同步性,并对复杂运动和真实环境动态具有鲁棒性。实验表明,相比现有配音流水线,本方法生成的视频在视觉保真度、口型同步性和鲁棒性方面均有提升。

原文摘要 · Abstract (English)

Audio-Visual Foundation Models, which are pretrained to jointly generate sound and visual content, have recently shown an unprecedented ability to model multi-modal generation and editing, opening new opportunities for downstream tasks. Among these tasks, video dubbing could greatly benefit from such priors, yet most existing solutions still rely on complex, task-specific pipelines that struggle in real-world settings. In this work, we introduce a single-model approach that adapts a foundational audio-video diffusion model for video-to-video dubbing via a lightweight LoRA. The LoRA enables the model to condition on an input audio-video while jointly generating translated audio and synchronized facial motion. To train this LoRA, we leverage the generative model itself to synthesize paired multilingual videos of the same speaker. Specifically, we generate multilingual videos with language switches within a single clip, and then inpaint the face and audio in each half to match the language of the other half. By leveraging the rich generative prior of the audio-visual model, our approach preserves speaker identity and lip synchronization while remaining robust to complex motion and real-world dynamics. We demonstrate that our approach produces high-quality dubbed videos with improved visual fidelity, lip synchronization, and robustness compared to existing dubbing pipelines.

视频配音扩散模型音视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。