arXiv:2510.03735cs.SD2025-10被引 2

通过频段分解实现音频编码器的解耦表征,提升可解释性与重建质量。

Soft Disentanglement in Frequency Bands for Neural Audio Codecs

  • 将时域信号做频段分解,多分支编码器分别处理各频段成分。
  • 重建和感知性能优于现有最优基线,且对修补任务有潜在优势。
  • 方法通用性强,不依赖特定数据或任务,适合音频解耦研究者。

在基于神经网络的音频特征提取中,确保表征捕捉解耦信息对模型可解释性至关重要。然而,现有解耦方法常依赖于高度依赖数据特性或特定任务的假设。本文提出一种可在神经架构中学习解耦特征的通用方法。该方法首先对时域信号进行谱分解,随后采用多分支音频编码器处理分解后的分量。实验评估表明,本方法在重建性能和感知质量上均优于当前最优基线,同时在音频修补任务中也展现出潜在优势。

原文摘要 · Abstract (English)

In neural-based audio feature extraction, ensuring that representations capture disentangled information is crucial for model interpretability. However, existing disentanglement methods often rely on assumptions that are highly dependent on data characteristics or specific tasks. In this work, we introduce a generalizable approach for learning disentangled features within a neural architecture. Our method applies spectral decomposition to time-domain signals, followed by a multi-branch audio codec that operates on the decomposed components. Empirical evaluations demonstrate that our approach achieves better reconstruction and perceptual performance compared to a state-of-the-art baseline while also offering potential advantages for inpainting tasks.

音频编码解耦表征频段分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。