用跨语料迁移学习提升语音障碍严重程度分类准确率
DSSCNet: A Transfer Learning Framework for Cross-Corpus Dysarthric Speech Severity Classification

- 通过多语料预训练+微调,增强模型泛化能力
- 在TORGO和UA-Speech数据集上分别达75.80%和68.25%准确率
- 适合开发自动化语音障碍评估工具的开发者
由于说话人差异、类别不平衡和数据量有限,语音障碍严重程度分类极具挑战。本文提出DSSCNet,一种基于迁移学习和多语料学习的深度学习框架,以提升无特定说话人依赖的分类性能。通过在一个语音障碍语料库上预训练,并在另一个上微调,DSSCNet显著提升了特征提取能力与跨语料泛化性。实验结果表明,该模型在无说话人依赖分类任务中优于现有最优模型,在TORGO数据集上达到75.80%准确率,在UA-Speech上达68.25%,大幅降低误分类率。研究证实,利用不同数据集间的知识迁移可增强模型鲁棒性,使DSSCNet适用于自动化语音障碍评估。本工作推动了面向言语障碍者的辅助语音技术发展。
原文摘要 · Abstract (English)
Dysarthric speech severity classification is challenging due to speaker variability, class imbalance, and limited datasets. This study introduces DSSCNet, a deep learning model that employs transfer learning and multi-corpus learning to enhance speaker-independent classification. By pre-training on one dysarthric speech corpus and fine-tuning on another, DSSCNet achieves improved feature extraction and cross-corpus generalization. Experimental results demonstrate that DSSCNet outperforms state-of-the-art models for speaker-independent severity classification, achieving 75.80\% accuracy on TORGO and 68.25\% on UA-Speech, significantly reducing misclassification errors. The findings confirm that leveraging knowledge transfer between datasets improves model robustness, making DSSCNet well-suited for automated dysarthria assessment. This research contributes to the development of more effective assistive speech technologies for individuals with speech impairments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。