arXiv:2510.09065cs.SDcs.CV2025-10中稿 · ICASSP 2026被引 4

基于预训练模型实现视频/文本查询的声音分离,高效且保持生成能力。

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation

  • 利用预训练视频到音频模型迁移知识,无需从头训练。
  • 在声音分离任务上优于现有确定性和生成类模型。
  • 微调后仍保留原始视频生成音频能力,适合多任务应用。

我们提出MMAudioSep,一种基于预训练视频到音频模型的生成式音效分离方法。通过利用预训练音频生成模型中学习到的视频/文本与音频之间的关联知识,该模型可高效训练,无需从零开始。我们在多个基准上对比了MMAudioSep与现有分离模型(包括确定性和生成式方法),结果表明其性能更优。此外,即使通过微调获得音效分离功能,模型仍保留原始视频到音频生成能力。这证明了基础音效生成模型在下游音效相关任务中的潜力。代码已开源。

原文摘要 · Abstract (English)

We introduce MMAudioSep, a generative model for video/text-queried sound separation that is founded on a pretrained video-to-audio model. By leveraging knowledge about the relationship between video/text and audio learned through a pretrained audio generative model, we can train the model more efficiently, i.e., the model does not need to be trained from scratch. We evaluate the performance of MMAudioSep by comparing it to existing separation models, including models based on both deterministic and generative approaches, and find it is superior to the baseline models. Furthermore, we demonstrate that even after acquiring functionality for sound separation via fine-tuning, the model retains the ability for original video-to-audio generation. This highlights the potential of foundational sound generation models to be adopted for sound-related downstream tasks. Our code is available at https://github.com/sony/mmaudiosep.

音效分离生成模型多模态迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。