arXiv:2606.09098eess.AS2026-06被引 1

一句话生成视频配音与音效,让配音更自然真实。

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis

论文配图:HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis
图 1 · 摘自论文原文
  • 用文本控制语音和音效联合生成,一次输入搞定
  • 在多个数据集上语音质量、同步性均超越现有方法
  • 适合影视配音、虚拟主播等需要复杂音效的场景

视频配音是多媒体内容创作的核心,旨在为视觉流生成同步的音频序列。尽管文本转语音(TTS)和文本转音频(TTA)已取得显著进展,现有配音系统仍局限于单一语音合成,未能融合音效和环境音,导致从业者依赖碎片化流程和繁琐的手动混音。为此,我们提出HoliDubber,一个突破语音单一生成的全景式视频配音框架,支持从单个文本提示中联合生成语音与音效。HoliDubber采用基于补丁的自回归扩散变压器架构,通过因果语言模型建模聚合补丁嵌入以捕捉全局时序结构,并利用扩散变压器解码器在每个补丁内生成高保真连续标记,遵循分而治之策略。为实现跨模态对齐,视觉特征被编码为补丁级表示,并通过交叉注意力与音频补丁融合,使语音生成能基于说话人面部动作动态进行。此外,我们构建了HoliDub-Bench,一个从既有数据集中筛选的包含同步视频-文本-音频三元组的基准数据集,用于全景配音评估。大量实验表明,HoliDubber在多个基准上显著优于现有方法,在语音质量、同步性和说话人相似性方面表现优异。HoliDub-Bench上的结果进一步验证了联合语音与音效生成的有效性,确立了复杂声学场景下全景视频配音的新范式。

原文摘要 · Abstract (English)

Video dubbing is a cornerstone of multimedia content creation, aiming to synthesize synchronized acoustic sequences for visual streams. While Text-to-Speech (TTS) and Text-to-Audio (TTA) generation have each achieved remarkable progress, existing dubbing systems remain confined to isolated speech synthesis without incorporating sound effects and ambient audio, forcing practitioners to rely on fragmented workflows and laborious manual post-mixing. To address this limitation, we present HoliDubber, a holistic video dubbing framework that moves beyond speech-only generation by enabling the joint synthesis of speech and sound effects from a single text prompt. Specifically, HoliDubber adopts a patch-based autoregressive diffusion transformer architecture, where a causal language model autoregressively models aggregated patch embeddings to capture global temporal structure, and a Diffusion Transformer decoder generates high-fidelity continuous tokens within each patch, following a divide-and-conquer strategy. To achieve cross-modal alignment, visual features are encoded into patch-level representations and fused with audio patches via cross-attention, enabling the model to ground speech generation in the speaker's visual articulation dynamics. In addition, we introduce HoliDub-Bench, a benchmark curated from established datasets with synchronized video-text-audio triplets designed for holistic dubbing evaluation. Extensive experiments demonstrate that HoliDubber significantly outperforms existing methods across multiple benchmarks in speech quality, synchronization, and speaker similarity. Furthermore, results on HoliDub-Bench validate the effectiveness of joint speech-and-sound generation, establishing a new paradigm for holistic video dubbing in complex acoustic scenes. \footnote{The demo page of the project is https://holidubber.github.io}

视频配音多模态生成音频合成跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。