arXiv:2509.01390cs.CLeess.AS2025-09被引 2

分析神经音频编码器的语义规律,发现其分词具语言特征。

Analysing the Language of Neural Audio Codecs

  • 对比多种模型的离散语音分词,研究其统计规律。
  • 3-gram分词符合齐普夫定律,且与语音识别准确率相关。
  • 为生成式语音模型设计提供语言结构依据。

本研究对神经音频编码器(NACs)产生的离散语音分词进行了统计与语言学特性对比分析。考察了不同NAC模型输出分词在齐普夫定律、希普斯定律、熵值与冗余度等方面的语言统计规律。通过自动语音识别错误率评估语音可懂度,以UTMOS分数评估语音质量,验证分词特性与语音保真度的关系。结果表明,尤其是3-gram分词序列表现出类语言统计特征;这些统计特性与信息量指标共同与语音识别和重建任务表现呈正相关。研究揭示了NAC分词序列的内在结构规律,为构建更高效的生成式语音模型提供理论支持。

原文摘要 · Abstract (English)

This study presents a comparative analysis of the statistical and linguistic properties of neural audio codecs (NACs). We investigate discrete speech tokens produced by various NAC models, examining their adherence to linguistic statistical laws such as Zipf's law and Heaps' law, as well as their entropy and redundancy. To assess how these token-level properties relate to semantic and acoustic preservation in synthesized speech, we evaluate intelligibility using error rates of automatic speech recognition, and quality using the UTMOS score. Our results reveal that NAC tokens, particularly 3-grams, exhibit language-like statistical patterns. Moreover, these properties, together with measures of information content, are found to correlate with improved performances in speech recognition and resynthesis tasks. These findings offer insights into the structure of NAC token sequences and inform the design of more effective generative speech models.

语音生成神经编码器语言规律统计分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。