从信息论角度评估语音离散化后的完整性和可访问性
Estimating the Completeness of Discrete Speech Units
- 用信息论方法分析残差向量量化前后信息完整性
- 发现HuBERT离散单元中说话人与音素信息均保留充分
- 建议挖掘残差中被丢弃的丰富信息,而非简单舍弃
使用离散语音单元在语音编码和生成中广泛存在,但关于自监督离散单元的若干假设尚未验证,如通过k-means实现音素与说话人信息解耦,或认为向量量化后存在信息损失。本文从信息论视角出发,评估量化前后的信息完整性(information completeness)与可访问性(information accessibility)。我们推导出信息完整性的下界,并对经过残差向量量化(Residual Vector Quantization)后的HuBERT表示进行完整性估计。结果表明,说话人信息在离散单元中仍充分存在,而音素信息则主要保留在残差中,说明向量量化并未实现有效解耦。研究为离散单元的选择提供了全面评估依据,提示应更多挖掘残差中的潜在信息,而非将其丢弃。
原文摘要 · Abstract (English)
Representing speech with discrete units has been widely used in speech codec and speech generation. However, there are several unverified claims about self-supervised discrete units, such as disentangling phonetic and speaker information with k-means, or assuming information loss after k-means. In this work, we take an information-theoretic perspective to answer how much information is present (information completeness) and how much information is accessible (information accessibility), before and after residual vector quantization. We show a lower bound for information completeness and estimate completeness on discretized HuBERT representations after residual vector quantization. We find that speaker information is sufficiently present in HuBERT discrete units, and that phonetic information is sufficiently present in the residual, showing that vector quantization does not achieve disentanglement. Our results offer a comprehensive assessment on the choice of discrete units, and suggest that a lot more information in the residual should be mined rather than discarded.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。