arXiv:2503.22200cs.SDcs.CV2025-03被引 1

通过分步思维引导提升视频-音频生成质量,效果显著优于现有模型。

Enhance Generation Quality of Flow Matching V2A Model via Multi-Step CoT-Like Guidance and Combined Preference Optimization

  • 采用类思维链的多阶段分步引导机制,实现专业级音效生成。
  • 在多个数据集上音质指标提升超40%,主观评分提高近20%。
  • 适合需要精准音视频同步的影视、游戏音效制作场景。

从视频和文本提示生成高质量音效需在语义与时间上精确对齐,并具备分步指导能力。现有最先进视频引导音效生成模型在通用与专业场景下仍表现不足。为此,本文提出端到端的多阶段、多模态生成框架Chain-of-Perform(CoP),采用基于Transformer的网络架构实现类思维链(CoT-like)引导,支持通用与专业音效生成。设计多阶段训练流程以确保高质量音效输出,并构建了由视频引导的CoP多模态数据集,支撑分步音效生成。评估结果表明,该框架在多个数据集上优于当前最优模型:VGGSound上FAD从0.79降至0.74(+6.33%),CLIP得分从16.12升至17.70(+9.80%);PianoYT-2h上SI-SDR从1.98dB提升至3.35dB(+69.19%),MOS从2.94升至3.49(+18.71%);Piano-10h上SI-SDR从2.22dB增至3.21dB(+44.59%),MOS从3.07升至3.42(+11.40%)。

原文摘要 · Abstract (English)

Creating high-quality sound effects from videos and text prompts requires precise alignment between visual and audio domains, both semantically and temporally, along with step-by-step guidance for professional audio generation. However, current state-of-the-art video-guided audio generation models often fall short of producing high-quality audio for both general and specialized use cases. To address this challenge, we introduce a multi-stage, multi-modal, end-to-end generative framework with Chain-of-Thought-like (CoT-like) guidance learning, termed Chain-of-Perform (CoP). First, we employ a transformer-based network architecture designed to achieve CoP guidance, enabling the generation of both general and professional audio. Second, we implement a multi-stage training framework that follows step-by-step guidance to ensure the generation of high-quality sound effects. Third, we develop a CoP multi-modal dataset, guided by video, to support step-by-step sound effects generation. Evaluation results highlight the advantages of the proposed multi-stage CoP generative framework compared to the state-of-the-art models on a variety of datasets, with FAD 0.79 to 0.74 (+6.33%), CLIP 16.12 to 17.70 (+9.80%) on VGGSound, SI-SDR 1.98dB to 3.35dB (+69.19%), MOS 2.94 to 3.49(+18.71%) on PianoYT-2h, and SI-SDR 2.22dB to 3.21dB (+44.59%), MOS 3.07 to 3.42 (+11.40%) on Piano-10h.

音视频生成分步引导生成质量多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。