arXiv:2409.04799cs.SDeess.AS2024-09中稿 · SLT 2024被引 2

用微调HuBERT构建发音特征原型,实现低资源构音障碍语音唤醒

PB-LRDWWS System for the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting Challenge

  • 通过三阶段微调HuBERT提取构音障碍语音特征,生成说话人专属原型
  • 在测试集上达到第二名,证明方法在低资源场景下有效
  • 适合研究语音识别、残障人士语音交互的开发者参考

针对SLT 2024低资源构音障碍语音唤醒挑战(LRDWWS),我们提出PB-LRDWWS系统。该系统结合构音障碍语音内容特征提取器与基于原型的分类方法。特征提取器为经过三阶段微调的HuBERT模型,采用交叉熵损失训练,从目标构音障碍说话人的注册语音中提取特征以构建原型。分类时,计算评估语音与原型之间的余弦相似度。尽管方法简单,实验结果表明其有效性。本系统在最终Test-B评测中获得第二名。

原文摘要 · Abstract (English)

For the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting (LRDWWS) Challenge, we introduce the PB-LRDWWS system. This system combines a dysarthric speech content feature extractor for prototype construction with a prototype-based classification method. The feature extractor is a fine-tuned HuBERT model obtained through a three-stage fine-tuning process using cross-entropy loss. This fine-tuned HuBERT extracts features from the target dysarthric speaker's enrollment speech to build prototypes. Classification is achieved by calculating the cosine similarity between the HuBERT features of the target dysarthric speaker's evaluation speech and prototypes. Despite its simplicity, our method demonstrates effectiveness through experimental results. Our system achieves second place in the final Test-B of the LRDWWS Challenge.

语音唤醒构音障碍HuBERT原型学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。