用视觉和文本联合控制,精准生成或替换视频音效。
AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
- 结合视觉、音频与文本语义,实现细粒度音效编辑。
- 在视频数据集上达到当前音效编辑最佳效果。
- 适合影视后期、音视频内容创作者使用。
音效编辑——通过添加、移除或替换元素来修改音频——仍受限于仅依赖低层信号处理或粗略文本提示的方法,常导致灵活性差且音频质量不佳。为此,我们提出 AV-Edit,一种基于视听语义联合控制的生成式音效编辑框架。该方法采用专门设计的对比音频-视觉掩码自编码器(CAV-MAE-Edit)进行多模态预训练,学习对齐的跨模态表征。这些表征用于训练一个编辑型多模态扩散变换器(MM-DiT),通过基于相关性的特征门控训练策略,实现移除视觉无关声音并生成与视频内容一致的缺失音频元素。此外,我们构建了一个专用的基于视频的音效编辑数据集作为评估基准。实验表明,所提方法能根据视频内容生成高质量音频,并实现精确修改,在音效编辑领域达到当前最优性能,并在音频生成领域展现出强竞争力。
原文摘要 · Abstract (English)
Sound effect editing-modifying audio by adding, removing, or replacing elements-remains constrained by existing approaches that rely solely on low-level signal processing or coarse text prompts, often resulting in limited flexibility and suboptimal audio quality. To address this, we propose AV-Edit, a generative sound effect editing framework that enables fine-grained editing of existing audio tracks in videos by jointly leveraging visual, audio, and text semantics. Specifically, the proposed method employs a specially designed contrastive audio-visual masking autoencoder (CAV-MAE-Edit) for multimodal pre-training, learning aligned cross-modal representations. These representations are then used to train an editorial Multimodal Diffusion Transformer (MM-DiT) capable of removing visually irrelevant sounds and generating missing audio elements consistent with video content through a correlation-based feature gating training strategy. Furthermore, we construct a dedicated video-based sound editing dataset as an evaluation benchmark. Experiments demonstrate that the proposed AV-Edit generates high-quality audio with precise modifications based on visual content, achieving state-of-the-art performance in the field of sound effect editing and exhibiting strong competitiveness in the domain of audio generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。