通过频谱分解提升神经音频编码器的可解释性,效果优于当前最佳方法。
Désentrelacement Fréquentiel Doux pour les Codecs Audio Neuronaux
- 利用时域信号的频谱分解实现特征解耦,增强表示可读性。
- 在重建保真度和主观听感上均超越现有最优基线。
- 适用于需要清晰语义结构的音频处理任务,如语音分离与合成。
尽管基于神经网络的模型在音频特征提取方面取得了显著进展,但学习到的表示可解释性仍是关键挑战。为解决此问题,已有研究将解耦技术引入离散神经音频编码器,以对提取的标记施加结构约束。然而,这些方法通常对特定数据集或任务设定依赖性强。本文提出一种基于频谱分解的解耦神经音频编码器,通过分析时域信号的频谱特性来提升表示可解释性。实验表明,该方法在重建保真度和感知质量上均优于当前最先进的基线模型。
原文摘要 · Abstract (English)
While neural-based models have led to significant advancements in audio feature extraction, the interpretability of the learned representations remains a critical challenge. To address this, disentanglement techniques have been integrated into discrete neural audio codecs to impose structure on the extracted tokens. However, these approaches often exhibit strong dependencies on specific datasets or task formulations. In this work, we propose a disentangled neural audio codec that leverages spectral decomposition of time-domain signals to enhance representation interpretability. Experimental evaluations demonstrate that our method surpasses a state-of-the-art baseline in both reconstruction fidelity and perceptual quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。