arXiv:2608.31037cs.CLcs.SD2026-08

分析13种神经音频编码器的离散令牌统计特性,揭示不同架构与噪声下的语言规律。

Language-Statistical Analysis of Neural Audio Codec Tokens Across Architectures, Corpora, and Noise Conditions

  • 对比多码本、单码本及非向量量化架构的令牌统计行为。
  • 白噪和真实噪声下,多码本编码器出现崩溃与爆炸退化现象。
  • 适用于研究语音编码器的统计特性或评估抗噪性能的开发者。

神经音频编码器(NACs)将语音转换为离散令牌序列,已有研究表明这些序列遵循类语言的统计规律。本文分析了13种NACs的令牌统计特性,涵盖多码本残差向量量化(RVQ)、单码本VQ及非VQ设计,基于三个语料库在干净、白噪声和真实世界DEMAND噪声条件下进行评估。通过匹配令牌样本并采用显式拟合有效性保障和族条件$ n $-gram阶数,估算Zipf参数、Heaps参数、一元熵、码本占用率及Jensen-Shannon散度(JSD)。语料身份对任何指标的方差解释力极低,而声学条件与量化器元类别在不同指标中主导影响,其中一元熵与元类别关联最强。在相同一元阶数下,干净到噪声的JSD在DEMAND噪声下与梅尔倒谱失真关联最明显。此前报道的RVQ编码器崩溃与爆炸退化现象分别集中于白噪声和DEMAND噪声下的RVQ单元;爆炸也出现在非VQ编码器中,而单码本VQ编码器仅表现出占用率与分布形态的变化,无上述两类退化迹象。结果为不同架构条件下的NAC令牌语言统计分析提供了规范依据。

原文摘要 · Abstract (English)

Neural audio codecs (NACs) convert speech into discrete token sequences, and prior work has reported that these sequences follow language-like statistical laws. This paper analyzes the token statistics of 13 NACs spanning multi-codebook residual vector quantization (RVQ), single-codebook VQ, and non-VQ designs, evaluated on three corpora under clean, white-noise, and real-world DEMAND-noise conditions. Zipf and Heaps parameters, unigram entropy, codebook occupancy, and Jensen-Shannon divergence (JSD) are estimated from matched token samples with explicit fit-validity safeguards and family-conditional $n$-gram orders. Corpus identity explains little variance in any metric, whereas acoustic condition and quantizer meta-category dominate in a metric-dependent way, and unigram entropy is the metric most strongly associated with meta-category. Clean-to-noise JSD computed at a common unigram order is associated with mel-cepstral distortion most clearly under DEMAND noise. The collapse and explosion degradation signatures previously reported for RVQ codecs concentrate in RVQ cells under white and DEMAND noise, respectively; explosion also occurs in non-VQ codecs, and single-codebook VQ codecs shift in occupancy and distribution shape without either signature. These results provide architecture-conditioned conventions for applying language-statistical analysis to NAC tokens.

音频编码语言统计向量量化噪声鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。