用大语言模型实现任意音频混合的通用分离,支持无限音源和场景。
UniSep: Universal Target Audio Separation with Language Models at Scale
- 将音频分离建模为序列生成任务,用大语言模型处理离散潜空间音频序列。
- 基于36.5小时大规模音频数据训练,覆盖语音、音乐与环境声,性能媲美专用模型。
- 提出纯音频预训练策略,减少仿真数据依赖,提升模型对音频内在关联的理解。
我们提出通用目标音频分离(UniSep),解决任意类型音频混合的分离问题。与以往研究不同,UniSep支持无限音源类别和数量。将分离任务建模为序列到序列问题,利用大语言模型(LLM)在离散潜空间中建模音频序列,借助其在大规模数据下的复杂混合音频处理能力。同时提出一种新型纯音频预训练策略,减少大规模数据仿真需求,增强LLM对音频序列内部一致性与相关性的理解。实验表明,在36.5小时大规模音频数据(含语音、音乐与声音)上训练的模型,可实现跨域通用分离,主观与客观评价均达到与单任务模型相当的性能。
原文摘要 · Abstract (English)
We propose Universal target audio Separation (UniSep), addressing the separation task on arbitrary mixtures of different types of audio. Distinguished from previous studies, UniSep is performed on unlimited source domains and unlimited source numbers. We formulate the separation task as a sequence-to-sequence problem, and a large language model (LLM) is used to model the audio sequence in the discrete latent space, leveraging the power of LLM in handling complex mixture audios with large-scale data. Moreover, a novel pre-training strategy is proposed to utilize audio-only data, which reduces the efforts of large-scale data simulation and enhances the ability of LLMs to understand the consistency and correlation of information within audio sequences. We also demonstrate the effectiveness of scaling datasets in an audio separation task: we use large-scale data (36.5k hours), including speech, music, and sound, to train a universal target audio separation model that is not limited to a specific domain. Experiments show that UniSep achieves competitive subjective and objective evaluation results compared with single-task models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。