arXiv:2507.10109cs.MMcs.SD2025-07被引 9

让视频自动生成同步的语音与背景音乐,突破传统只生成背景音的局限。

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis

  • 用统一框架同时生成语音和背景音,通过跨模态对齐提升同步性。
  • 在新构建的DualBench基准上,生成音轨质量高且时间对齐精准。
  • 适合做视频自动配乐、影视后期合成的研究者与开发者使用。

尽管近期视频到音频(V2A)模型能从视觉输入生成逼真的背景音,但普遍忽略语音——这正是许多视频音轨的核心部分。本文提出新任务:视频到音轨(V2ST)生成,旨在统一框架中联合生成同步的背景音与语音。为此,我们提出DualDub,一个基于多模态语言模型的统一框架,包含多模态编码器、跨模态对齐器以及双解码头,实现语音与背景音的并行生成。特别地,跨模态对齐器采用因果与非因果注意力机制,提升音轨同步性与声学和谐度。针对数据稀缺问题,设计渐进式课程学习策略以逐步增强多模态能力。最后,构建DualBench——首个用于V2ST评估的基准,包含精心筛选的测试集与全面指标。实验表明,DualDub达到当前最佳性能,生成高质量且高度同步的音轨。

原文摘要 · Abstract (English)

While recent video-to-audio (V2A) models can generate realistic background audio from visual input, they largely overlook speech, an essential part of many video soundtracks. This paper proposes a new task, video-to-soundtrack (V2ST) generation, which aims to jointly produce synchronized background audio and speech within a unified framework. To tackle V2ST, we introduce DualDub, a unified framework built on a multimodal language model that integrates a multimodal encoder, a cross-modal aligner, and dual decoding heads for simultaneous background audio and speech generation. Specifically, our proposed cross-modal aligner employs causal and non-causal attention mechanisms to improve synchronization and acoustic harmony. Besides, to handle data scarcity, we design a curriculum learning strategy that progressively builds the multimodal capability. Finally, we introduce DualBench, the first benchmark for V2ST evaluation with a carefully curated test set and comprehensive metrics. Experimental results demonstrate that DualDub achieves state-of-the-art performance, generating high-quality and well-synchronized soundtracks with both speech and background audio.

音轨生成多模态语音合成视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。