arXiv:2511.16126eess.ASeess.SP2025-11被引 1

让音频编码区分不同声源,可按需选择编码特定声音

SUNAC: Source-aware Unified Neural Audio Codec

  • 根据声源类型提示,从混音中直接分离并编码单个声源
  • 在保持高保真度的同时,计算成本低于先分离后编码的流程
  • 适合需要灵活提取特定声音的应用,如语音识别或声音分析

神经音频编解码器(NAC)能生成紧凑表示,适用于大语言模型等下游任务。然而大多数NAC对多个声源的混合信号进行纠缠编码,不利于仅需部分声源的应用(如特定说话人转录)。为此,我们提出一种源感知编解码器,可基于声源类型提示,直接从混音中编码单个声源,支持用户按需选择编码目标,包括同类型多个声源的独立编码。实验表明,该模型在重建和分离质量上媲美先分离后编码的级联方案,且计算开销更低。

原文摘要 · Abstract (English)

Neural audio codecs (NACs) provide compact representations that can be leveraged in many downstream applications, in particular large language models. Yet most NACs encode mixtures of multiple sources in an entangled manner, which may impede efficient downstream processing in applications that need access to only a subset of the sources (e.g., analysis of a particular type of sound, transcription of a given speaker, etc). To address this, we propose a source-aware codec that encodes individual sources directly from mixtures, conditioned on source type prompts. This enables user-driven selection of which source(s) to encode, including separately encoding multiple sources of the same type (e.g., multiple speech signals). Experiments show that our model achieves competitive resynthesis and separation quality relative to a cascade of source separation followed by a conventional NAC, with lower computational cost.

音频编码声源分离神经编解码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。