探测神经音频编码器能否捕捉英语语调差异,发现其对语调分类有潜力但仍有不足。
Probing neural audio codecs for distinctions among English nuclear tunes
- 用线性探针分析编码器潜在表示,测试其是否包含语调信息
- 对五类核心语调的分类准确率达45%,二元区分可达89%
- 揭示语调信息分散在多个代码本中,挑战了语义/声学分离的传统观点
当前最先进的语音对话模型(Défossez et al. 2024;Schalkwyk et al. 2025)使用神经音频编码器将音频信号转换为低频向量潜表示序列,每个向量通过多层向量码本进行量化。变压器层使这些表示能够反映时间与上下文相关的模式。我们在Cole等(2023)提供的标注音频数据上训练探针,检验表征英语句末(核)语调的音高轨迹是否属于这些模式。结果显示:在线性探针上,基于未量化潜表示或部分关联码字的分类,对八种具有单音调重音的音调类型(最高平均测试准确率,TATA:0.31),以及在人类语音产生与感知中稳健的五类音调聚类(TATA:0.45)均达到高于随机水平的准确率。在上升与下降语调两类的二元区分上,准确率更高(TATAs:0.74–0.89),分别用于疑问句与陈述句。语调信息分布在所有码本中,质疑了文献中‘语义’与‘声学’码本的区分。非线性探针进一步提升性能,但对五类聚类的判别仍远低于人类水平,表明当前编码器存在根本性局限。
原文摘要 · Abstract (English)
State-of-the-art spoken dialogue models (Défossez et al. 2024; Schalkwyk et al. 2025) use neural audio codecs to "tokenize" audio signals into a lower-frequency stream of vectorial latent representations, each quantized using a hierarchy of vector codebooks. A transformer layer allows these representations to reflect some time- and context-dependent patterns. We train probes on labeled audio data from Cole et al. (2023) to test whether the pitch trajectories that characterize English phrase-final (nuclear) intonational tunes are among these patterns. Results: Linear probes trained on the unquantized latents or some of the associated codewords yield above-chance accuracy in distinguishing eight phonologically specified nuclear tunes with monotonal pitch accents (top average test accuracy (TATA): 0.31) and the five clusters of these tunes that are robust in human speech production and perception (TATA: 0.45). Greater accuracy (TATAs: 0.74-0.89) is attained for binary distinctions between classes of rising vs. falling tunes, respectively used for questions and assertions. Information about tunes is spread among all codebooks, which calls into question a distinction between 'semantic' and 'acoustic' codebooks found in the literature. Accuracies improve with nonlinear probes, but discrimination among the five clusters remains far from human performance, suggesting a fundamental limitation of current codecs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。