arXiv:2511.04623cs.SDeess.AS2025-11被引 4

用语音模仿+文本提示,实现更灵活的音频分离与去除

PromptSep: Generative Audio Separation via Multimodal Prompting

  • 结合文本和人声模仿作为输入条件,增强模型控制能力
  • 在多个数据集上实现音源分离与声音移除的顶尖效果
  • 适合需要精细音频编辑的创作者或音效工程师

近期语言查询音频源分离(LASS)取得突破,表明生成模型可实现优于传统掩码方法的音频分离质量。但其实际应用受限于两点:一是用户常需除分离外的其他操作,如声音移除;二是仅靠文本提示难以直观指定音源。本文提出PromptSep,将LASS扩展为通用音频分离框架。该方法采用条件扩散模型,并通过精细化数据模拟提升性能,支持音频提取与声音移除。为突破纯文本查询的局限,引入人声模仿作为更直观的条件模态,结合Sketch2Sound策略进行数据增强。在多个基准上的客观与主观评估均表明,PromptSep在声音移除和基于人声模仿的源分离任务中达到当前最优表现,同时保持在语言查询源分离上的竞争力。

原文摘要 · Abstract (English)

Recent breakthroughs in language-queried audio source separation (LASS) have shown that generative models can achieve higher separation audio quality than traditional masking-based approaches. However, two key limitations restrict their practical use: (1) users often require operations beyond separation, such as sound removal; and (2) relying solely on text prompts can be unintuitive for specifying sound sources. In this paper, we propose PromptSep to extend LASS into a broader framework for general-purpose sound separation. PromptSep leverages a conditional diffusion model enhanced with elaborated data simulation to enable both audio extraction and sound removal. To move beyond text-only queries, we incorporate vocal imitation as an additional and more intuitive conditioning modality for our model, by incorporating Sketch2Sound as a data augmentation strategy. Both objective and subjective evaluations on multiple benchmarks demonstrate that PromptSep achieves state-of-the-art performance in sound removal and vocal-imitation-guided source separation, while maintaining competitive results on language-queried source separation.

音频分离多模态扩散模型语音模仿

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。