统一多模态声音分离,支持文本、图像、音频查询。
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
- 用查询混合法融合多模态特征,统一训练框架。
- 在三个数据集上达到当前最佳性能,支持开集分离。
- 可正负调控声音保留或移除,适合多媒体应用。
近年来,视觉与语言领域通过规模扩展取得了巨大成功。然而,在音频领域,研究者面临训练数据难以扩大的挑战,因为自然音频中常包含多种干扰信号。为此,我们提出统一多模态声音分离(OmniSep)框架,能够基于单模态或多模态组合查询分离出纯净音轨。具体而言,引入查询混合法(Query-Mixup),在训练中混合不同模态的查询特征,使OmniSep能同时优化多模态输入,将所有模态纳入统一框架。此外,允许查询对分离结果产生正向或负向影响,实现对特定声音的保留或剔除。最后,OmniSep采用检索增强策略Query-Aug,支持开词汇声音分离。在MUSIC、VGGSOUND-CLEAN+和MUSIC-CLEAN+数据集上的实验表明,该方法在文本、图像、音频查询任务中均达到领先性能。更多样例与信息请访问演示页:https://omnisep.github.io/。
原文摘要 · Abstract (English)
The scaling up has brought tremendous success in the fields of vision and language in recent years. When it comes to audio, however, researchers encounter a major challenge in scaling up the training data, as most natural audio contains diverse interfering signals. To address this limitation, we introduce Omni-modal Sound Separation (OmniSep), a novel framework capable of isolating clean soundtracks based on omni-modal queries, encompassing both single-modal and multi-modal composed queries. Specifically, we introduce the Query-Mixup strategy, which blends query features from different modalities during training. This enables OmniSep to optimize multiple modalities concurrently, effectively bringing all modalities under a unified framework for sound separation. We further enhance this flexibility by allowing queries to influence sound separation positively or negatively, facilitating the retention or removal of specific sounds as desired. Finally, OmniSep employs a retrieval-augmented approach known as Query-Aug, which enables open-vocabulary sound separation. Experimental evaluations on MUSIC, VGGSOUND-CLEAN+, and MUSIC-CLEAN+ datasets demonstrate effectiveness of OmniSep, achieving state-of-the-art performance in text-, image-, and audio-queried sound separation tasks. For samples and further information, please visit the demo page at \url{https://omnisep.github.io/}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。