arXiv:2506.04492eess.AS2025-06中稿 · ed被引 6

让神经音频编码器的输出变得可解释,直接从编码单元提取语音特征。

Bringing Interpretability to Neural Audio Codecs

  • 分两步分析:先理解语音信息如何编码,再用新网络解码特征。
  • 能从编码单元直接提取内容、身份和音高等语音属性。
  • 适合研究音频编码可解释性或想提取语音特征的研究者。

神经音频编码器因其在使用变换器高效建模音频方面的潜力而日益流行。这类先进编码器将连续波形转换为低采样离散单元。与语义单元不同,声学单元因训练目标主要聚焦于重建性能,往往缺乏可解释性。本文提出一种两阶段方法,探索编码器中语音信息的编码方式。分析阶段旨在深入理解内容、身份和音高等语音属性如何被编码;合成阶段则训练一个AnCoGen网络,对编码器进行后处理解释,直接从相应编码单元中提取语音属性。

原文摘要 · Abstract (English)

The advent of neural audio codecs has increased in popularity due to their potential for efficiently modeling audio with transformers. Such advanced codecs represent audio from a highly continuous waveform to low-sampled discrete units. In contrast to semantic units, acoustic units may lack interpretability because their training objectives primarily focus on reconstruction performance. This paper proposes a two-step approach to explore the encoding of speech information within the codec tokens. The primary goal of the analysis stage is to gain deeper insight into how speech attributes such as content, identity, and pitch are encoded. The synthesis stage then trains an AnCoGen network for post-hoc explanation of codecs to extract speech attributes from the respective tokens directly.

音频编码可解释性语音特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。