arXiv:2509.11717cs.SDcs.LG2025-09中稿 · Transactions on Ma…被引 1

在音频编码隐空间直接分离声音,实现低延迟通用音源分割。

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents

  • 用提示词驱动的轻量 Transformer 在编码隐空间直接分离音源
  • 端到端仅需1.35 GMAC,比AudioSep少54倍计算量
  • 支持开放词汇、低延迟部署,适合边缘设备实时处理

文本引导的声音分离可实现灵活音频编辑与开放域音源提取,但如AudioSep等系统因计算开销大,难以在低延迟边缘或编码器传输场景部署。现有神经音频编码器分离方法虽高效,但多局限于固定音源或封闭分类体系。本文提出CodecSep,一种在神经音频编码隐空间直接进行提示驱动的通用声音分离框架。该框架结合冻结的DAC骨干网络与轻量级FiLM调制Transformer掩码器,由CLAP文本嵌入驱动,实现开放词汇分离并保持编码器原生效率。在dnr-v2及五个开放域基准上,CodecSep在SI-SDR指标上持续优于AudioSep,ViSQOL表现相当,人类评分(MOS-LQS)显著提升。控制分析表明,细粒度提示优于粗略标签,显式隐空间掩码远胜于解码器风格生成。定性诊断显示,神经音频编码隐空间保留了源相关的结构特征,CodecSep主要通过通道级源条件调制加以利用。此外,该框架提供实际编码流部署路径:音频以编码码流传输时,CodecSep将码流映射为嵌入,在编码空间中直接分离,输出波形或重新量化码流,避免解码-分离-重编码循环。在此模式下,端到端仅需1.35 GMAC,相较AudioSep同流程节省约54倍算力,分离模块仅需25倍更低算力,且延迟和内存占用大幅降低。更广泛而言,CodecSep为编码器原生下游音频处理提供了范例。

原文摘要 · Abstract (English)

Text-guided sound separation enables flexible audio editing, assistive listening, and open-domain source extraction, but systems such as AudioSep remain too expensive for low-latency edge or codec-mediated deployment. Existing neural audio codec separators are efficient, yet largely restricted to fixed stems or closed taxonomies. We introduce CodecSep, a prompt-driven universal sound separation framework that extracts sources directly in neural audio codec latent space. CodecSep combines a frozen DAC backbone with a lightweight FiLM-conditioned Transformer masker driven by CLAP text embeddings, enabling open-vocabulary separation while preserving codec-native efficiency. Across dnr-v2 and five open-domain benchmarks, CodecSep consistently improves over AudioSep in SI-SDR, remains competitive in ViSQOL, and achieves clear gains in human MOS-LQS. Controlled analyses show that fine-grained prompts outperform coarse labels, and that explicit latent masking is substantially more effective than decoder-style latent generation in codec space. Qualitative diagnostics show that neural audio codec latents retain source-dependent structure, which CodecSep exploits mainly through channel-wise source-conditioned modulation. CodecSep also provides a practical code-stream deployment path. When audio is transmitted as neural audio codec codes, CodecSep maps codes to embeddings, separates directly in codec space, and outputs waveforms or re-quantized codes, avoiding the decode-separate-re-encode loop. In this regime, CodecSep requires only 1.35 GMACs end-to-end: about 54 times less compute than AudioSep in the same pipeline and 25 times lower separator-only compute, with much lower latency and memory. More broadly, CodecSep offers a blueprint for codec-native downstream audio processing.

声音分离编码器低延迟提示驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。