arXiv:2409.12528cs.SDeess.AS2024-09被引 6

用音频基础模型M2D增强SoundBeam,提升目标声音提取效果。

SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model

  • 将M2D音频特征融入SoundBeam系统
  • 使用录音查询时性能显著提升
  • 适用于多种声音类型,实用性强

目标声音提取(TSE)旨在通过线索从混响声中分离出特定声音。该任务需同时解决目标源识别与信号提取问题,且要求系统适应各类声音。为应对训练困难,本文引入预训练的音频基础模型M2D,其通过双重目标(声音标签预测与掩码预测)学习声音特征,与TSE任务高度契合。我们提出一种新TSE系统,将M2D的特征表示集成到SoundBeam中,该系统可利用声音类别标签或预录语音查询作为线索。实验表明,引入M2D能有效提升提取性能,尤其在使用语音查询时优势明显。

原文摘要 · Abstract (English)

Target sound extraction (TSE) consists of isolating a desired sound from a mixture of arbitrary sounds using clues to identify it. A TSE system requires solving two problems at once, identifying the target source and extracting the target signal from the mixture. For increased practicability, the same system should work with various types of sound. The duality of the problem and the wide variety of sounds make it challenging to train a powerful TSE system from scratch. In this paper, to tackle this problem, we explore using a pre-trained audio foundation model that can provide rich feature representations of sounds within a TSE system. We chose the masked-modeling duo (M2D) foundation model, which appears especially suited for the TSE task, as it is trained using a dual objective consisting of sound-label predictions and improved masked prediction. These objectives are related to sound identification and the signal extraction problems of TSE. We propose a new TSE system that integrates the feature representation from M2D into SoundBeam, which is a strong TSE system that can exploit both target sound class labels and pre-recorded enrollments (or audio queries) as clues. We show experimentally that using M2D can increase extraction performance, especially when employing enrollment clues.

声音提取音频模型M2DSoundBeam

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。