arXiv:2503.22208cs.SDcs.CV2025-03被引 7

用大模型的思维链生成视频配乐,让声音更准更自然。

DeepSound-V1: Start to Think Step-by-Step in the Audio Generation from Videos

  • 让多模态大模型逐步推理,不依赖额外标注
  • 声音错位率降低超38%,音质和同步性提升显著
  • 适合做视频配音、影视后期的开发者参考

目前高质量、同步的音频生成主要依赖多模态联合学习框架,但视觉与音频之间的对齐仍不理想。主要原因在于开源视频-音频和文本-音频数据集中缺乏足够的时序与语义对齐标注。为此,我们提出一种基于多模态大语言模型内部思维链(CoT)的音频生成框架,实现无需额外标注的分步推理。同时构建了一个多模态推理数据集,支持音频生成的初始推理学习。实验表明,该方法显著减少生成音频的错位(如语音重叠),性能优于多种先进模型。具体指标显示,F DP aSST 下降最多10.07%,F DP AN N s 下降最多11.62%,F DV GG 下降最多38.61%;IS 提升4.95%,IB-score 增加6.39%,DeSync 降低0.89%。

原文摘要 · Abstract (English)

Currently, high-quality, synchronized audio is synthesized from video and optional text inputs using various multi-modal joint learning frameworks. However, the precise alignment between the visual and generated audio domains remains far from satisfactory. One key factor is the lack of sufficient temporal and semantic alignment annotations in open-source video-audio and text-audio benchmarks. Therefore, we propose a framework for audio generation from videos, leveraging the internal chain-of-thought (CoT) of a multi-modal large language model (MLLM) to enable step-by-step reasoning without requiring additional annotations. Additionally, a corresponding multi-modal reasoning dataset is constructed to facilitate the learning of initial reasoning in audio generation. In the experiments, we demonstrate the effectiveness of the proposed framework in reducing misalignment (voice-over) in generated audio and achieving competitive performance compared to various state-of-the-art models. The evaluation results show that the proposed method outperforms state-of-the-art approaches across multiple metrics. Specifically, the F DP aSST indicator is reduced by up to 10.07%, the F DP AN N s indicator by up to 11.62%, and the F DV GG indicator by up to 38.61%. Furthermore, the IS indicator improves by up to 4.95%, the IB-score indicator increases by up to 6.39%, and the DeSync indicator is reduced by up to 0.89%.

音频生成思维链多模态视频配乐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。