arXiv:2410.23987eess.AScs.SD2024-10被引 10

一个模型搞定五类音频分离任务,靠提示词灵活切换模式。

Task-Aware Unified Source Separation

  • 用可学习的提示词控制分离行为,动态适应不同任务
  • 在5个主要音频分离任务上表现良好,包括矛盾需求场景
  • 适合需要统一处理多类型音频分离的工业应用

针对语音增强、语音分离、声事件分离、音乐源分离(MSS)及电影音频分离(CASS)等多任务分离问题,现有模型常因任务间目标冲突(如MSS要求分离乐器,而CASS需合并乐器)难以兼顾。为此,本文提出任务感知的统一源分离(TUSS)模型,通过可学习的变量提示词指定目标源,使模型能根据输入提示动态调整分离策略,从而统一处理所有主流分离任务。实验表明,该模型成功实现了上述五类任务的联合处理。我们还提供了合成与真实录音的音频样例,展示其在推理阶段依据提示灵活改变行为的能力。

原文摘要 · Abstract (English)

Several attempts have been made to handle multiple source separation tasks such as speech enhancement, speech separation, sound event separation, music source separation (MSS), or cinematic audio source separation (CASS) with a single model. These models are trained on large-scale data including speech, instruments, or sound events and can often successfully separate a wide range of sources. However, it is still challenging for such models to cover all separation tasks because some of them are contradictory (e.g., musical instruments are separated in MSS while they have to be grouped in CASS). To overcome this issue and support all the major separation tasks, we propose a task-aware unified source separation (TUSS) model. The model uses a variable number of learnable prompts to specify which source to separate, and changes its behavior depending on the given prompts, enabling it to handle all the major separation tasks including contradictory ones. Experimental results demonstrate that the proposed TUSS model successfully handles the five major separation tasks mentioned earlier. We also provide some audio examples, including both synthetic mixtures and real recordings, to demonstrate how flexibly the TUSS model changes its behavior at inference depending on the prompts.

音频分离多任务提示词统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。