arXiv:2603.18359cs.SD2026-03

用稀疏自编码器揭示语音编码器中的口音信息编码规律

Towards Interpretable Framework for Neural Audio Codecs via Sparse Autoencoders: A Case Study on Accent Information

  • 用稀疏自编码器分解神经语音编码器的密集表示,使其可解释
  • 语音导向模型靠激活强度编码口音,音位导向模型依赖激活位置
  • 低比特率EnCodec模型解释性更高,适合敏感场景应用

神经语音编码器(NACs)在现代语音系统中广泛应用,但其如何编码语言与副语言信息仍不清晰。提升NAC表示的可解释性对理解与部署于敏感场景至关重要。为此,我们采用稀疏自编码器(SAEs)将密集的NAC表示分解为稀疏且可解释的激活。本文聚焦具有挑战性的副语言属性——口音,提出量化NAC可解释性的框架。在16种SAE配置下评估四种NAC模型,使用相对性能指数进行分析。结果表明,DAC与SpeechTokenizer具备最高可解释性。进一步发现:声学导向型NAC主要通过稀疏表示的激活幅度编码口音信息,而音位导向型则更依赖激活位置;低比特率的EnCodec变体表现出更高的可解释性。

原文摘要 · Abstract (English)

Neural Audio Codecs (NACs) are widely adopted in modern speech systems, yet how they encode linguistic and paralinguistic information remains unclear. Improving the interpretability of NAC representations is critical for understanding and deploying them in sensitive applications. Hence, we employ Sparse Autoencoders (SAEs) to decompose dense NAC representations into sparse, interpretable activations. In this work, we focus on a challenging paralinguistic attribute-accent-and propose a framework to quantify NAC interpretability. We evaluate four NAC models under 16 SAE configurations using a relative performance index. Our results show that DAC and SpeechTokenizer achieve the highest interpretability. We further reveal that acoustic-oriented NACs encode accent information primarily in activation magnitudes of sparse representations, whereas phonetic-oriented NACs rely more on activation positions, and that low-bitrate EnCodec variants show higher interpretability.

语音编码可解释性稀疏编码口音识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。