arXiv:2412.09168cs.SDcs.CV2024-12被引 6

用视频生成高质量音效,少样本也能搞定

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls

  • 通过多模态思维链控制,让音视频精准对齐
  • 在少样本条件下生成高保真同步音效
  • 适合影视、游戏等需要快速音效制作的场景

为产品级视频生成音效时,真实场景下标注数据极少,需在少样本条件下生成高质量声音。为此,我们提出YingSound——一种面向视频引导音效生成的基础模型,支持少样本高质量音频生成。该模型包含两个核心模块:第一模块采用条件流匹配变压器,实现音视频语义对齐,构建可学习的音视频聚合器(AVA),在多阶段融合高分辨率视觉特征与对应音频特征;第二模块引入多模态视觉-音频思维链(CoT)方法,在少样本设置下生成更精细的音效。此外,我们构建了一个涵盖多种真实场景的行业标准视频到音频(V2A)数据集。实验表明,通过自动化评估和人工测试,YingSound能在多样条件输入下有效生成高质量且同步的音效。

原文摘要 · Abstract (English)

Generating sound effects for product-level videos, where only a small amount of labeled data is available for diverse scenes, requires the production of high-quality sounds in few-shot settings. To tackle the challenge of limited labeled data in real-world scenes, we introduce YingSound, a foundation model designed for video-guided sound generation that supports high-quality audio generation in few-shot settings. Specifically, YingSound consists of two major modules. The first module uses a conditional flow matching transformer to achieve effective semantic alignment in sound generation across audio and visual modalities. This module aims to build a learnable audio-visual aggregator (AVA) that integrates high-resolution visual features with corresponding audio features at multiple stages. The second module is developed with a proposed multi-modal visual-audio chain-of-thought (CoT) approach to generate finer sound effects in few-shot settings. Finally, an industry-standard video-to-audio (V2A) dataset that encompasses various real-world scenarios is presented. We show that YingSound effectively generates high-quality synchronized sounds across diverse conditional inputs through automated evaluations and human studies. Project Page: \url{https://giantailab.github.io/yingsound/}

音效生成多模态少样本视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。