根据文字提示,从多物体视频中只生成目标声音。
Hear What Matters! Text-conditioned Selective Video-to-Audio Generation
- 用文本作为选择器,提取视频中相关声源的视觉特征。
- 在真实数据集上音频质量、语义对齐和时间同步均表现优异。
- 适合需要精准音频编辑与创作的多媒体制作场景。
本文提出一种新任务——文本条件下的选择性视频到音频(V2A)生成,旨在从包含多个物体的视频中仅生成用户指定的声音。该能力在多媒体制作中至关重要,可实现对每个音源的独立音频处理,便于精确编辑、混音与创意控制。为此,我们提出SELVA模型,将文本提示视为显式选择器,从视频编码器中区分并提取与提示相关的声源视觉特征。为高效微调视频编码器并抑制无关激活,引入辅助标记以促进跨注意力机制,实现稳健的语义与时间定位。此外,采用自监督方式的自主视频混合方案,克服单音轨监督数据不足的问题。我们在VGG-MONOAUDIO数据集上进行评估,该数据集是专为该任务构建的高质量单音源视频基准。大量实验与消融分析一致验证了模型在音频质量、语义对齐和时间同步方面的有效性。
原文摘要 · Abstract (English)
This work introduces a new task, text-conditioned selective video-to-audio (V2A) generation, which produces only the user-intended sound from a multi-object video. This capability is especially crucial in multimedia production, where audio tracks are handled individually for each sound source for precise editing, mixing, and creative control. We propose SELVA, a novel text-conditioned V2A model that treats the text prompt as an explicit selector to distinctly extract prompt-relevant sound-source visual features from the video encoder. To suppress text-irrelevant activations with efficient video encoder finetuning, the proposed supplementary tokens promote cross-attention to yield robust semantic and temporal grounding. SELVA further employs an autonomous video-mixing scheme in a self-supervised manner to overcome the lack of mono audio track supervision. We evaluate SELVA on VGG-MONOAUDIO, a curated benchmark of clean single-source videos for such a task. Extensive experiments and ablations consistently verify its effectiveness across audio quality, semantic alignment, and temporal synchronization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。