arXiv:2503.22265cs.CVcs.SD2025-03被引 7

端到端生成视频与文本驱动的同步语音和音频,提升配音真实感。

DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation

  • 融合视频与文本输入,分阶段生成语音与背景音。
  • 在语音生成任务中,字错率降至3.15%,音色相似度达89.38%。
  • 适合影视配音、虚拟人语音合成等多模态场景应用。

当前高质量同步音频生成依赖于多模态联合学习框架,利用视频与可选文本输入。在视频转音频基准测试中,语音质量、语义对齐与音画同步已取得良好效果。然而在真实场景中,语音与音频常同时存在,基于视频与文本条件的端到端同步语音与音频生成仍缺乏研究。为此,我们提出一种端到端多模态生成框架DeepAudio,可同时根据视频与文本生成语音与音频。该框架包含视频转音频(V2A)、文本转语音(TTS)及动态模态融合(MoF)模块。在多个基准测试中,该框架达到领先水平:在视频-语音任务中,字错率(WER)从16.57%降至3.15%(+80.99%),音色相似度(SPK-SIM)从78.30%升至89.38%(+14.15%),情感相似度(EMO-SIM)从66.24%升至75.56%(+14.07%),梅尔倒谱失真(MCD)从8.59降至7.98(+7.10%),短时平均谱失真(MCD SL)从11.05降至9.40(+14.93%),适用于多种配音设置。

原文摘要 · Abstract (English)

Currently, high-quality, synchronized audio is synthesized using various multi-modal joint learning frameworks, leveraging video and optional text inputs. In the video-to-audio benchmarks, video-to-audio quality, semantic alignment, and audio-visual synchronization are effectively achieved. However, in real-world scenarios, speech and audio often coexist in videos simultaneously, and the end-to-end generation of synchronous speech and audio given video and text conditions are not well studied. Therefore, we propose an end-to-end multi-modal generation framework that simultaneously produces speech and audio based on video and text conditions. Furthermore, the advantages of video-to-audio (V2A) models for generating speech from videos remain unclear. The proposed framework, DeepAudio, consists of a video-to-audio (V2A) module, a text-to-speech (TTS) module, and a dynamic mixture of modality fusion (MoF) module. In the evaluation, the proposed end-to-end framework achieves state-of-the-art performance on the video-audio benchmark, video-speech benchmark, and text-speech benchmark. In detail, our framework achieves comparable results in the comparison with state-of-the-art models for the video-audio and text-speech benchmarks, and surpassing state-of-the-art models in the video-speech benchmark, with WER 16.57% to 3.15% (+80.99%), SPK-SIM 78.30% to 89.38% (+14.15%), EMO-SIM 66.24% to 75.56% (+14.07%), MCD 8.59 to 7.98 (+7.10%), MCD SL 11.05 to 9.40 (+14.93%) across a variety of dubbing settings.

语音生成多模态端到端配音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。