arXiv:2501.04379cs.SDeess.AS2025-01被引 5

用语音纯度引导生成更准的离散音素特征,提升失语症语音识别效果。

Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition

  • 引入语音纯度监督优化聚类与向量量化,增强音素区分能力
  • 在UASpeech数据集上相对降低4.82%错误率,最低达23.25% WER
  • 适合研究语音识别中失语症建模与离散表征优化的学者

从连续语音特征中提取的离散令牌能提供高效且领域适应的语音表征,但其在发音不准确、与正常语音差异大的失语症语音中的应用尚未探索。为提升无监督K-means或向量量化过程中因音素判别力减弱导致的性能问题,本文提出一种新的语音纯度引导(PPG)离散令牌方法。通过音素标签监督正则化标准K-means和基于变分自编码器-向量量化(VAE-VQ)的离散令牌提取中的最大似然与重构误差。在包含16名失语症患者的UASpeech语料库上实验表明,使用HuBERT提取的PPG离散令牌,在不同码本大小下均显著优于非PPG的K-means或VAE-VQ令牌,分别带来0.99%和1.77%的绝对词错误率(WER)降低(相对降低3.21%和4.82%)。结合不同令牌特征的系统获得最低WER 23.25%。在语音纯度指标上也取得一致提升。T-SNE可视化显示,引入语音纯度引导后,K-means/VAE-VQ聚类间的决策边界更加清晰。

原文摘要 · Abstract (English)

Discrete tokens extracted provide efficient and domain adaptable speech features. Their application to disordered speech that exhibits articulation imprecision and large mismatch against normal voice remains unexplored. To improve their phonetic discrimination that is weakened during unsupervised K-means or vector quantization of continuous features, this paper proposes novel phone-purity guided (PPG) discrete tokens for dysarthric speech recognition. Phonetic label supervision is used to regularize maximum likelihood and reconstruction error costs used in standard K-means and VAE-VQ based discrete token extraction. Experiments conducted on the UASpeech corpus suggest that the proposed PPG discrete token features extracted from HuBERT consistently outperform hybrid TDNN and End-to-End (E2E) Conformer systems using non-PPG based K-means or VAE-VQ tokens across varying codebook sizes by statistically significant word error rate (WER) reductions up to 0.99\% and 1.77\% absolute (3.21\% and 4.82\% relative) respectively on the UASpeech test set of 16 dysarthric speakers. The lowest WER of 23.25\% was obtained by combining systems using different token features. Consistent improvements on the phone purity metric were also achieved. T-SNE visualization further demonstrates sharper decision boundaries were produced between K-means/VAE-VQ clusters after introducing phone-purity guidance.

语音识别失语症离散表征音素纯度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。