提出DSSCNet模型与跨语料微调方法,提升失语症语音严重程度分类准确率。
Enhancing Speaker-Independent Dysarthric Speech Severity Classification with DSSCNet and Cross-Corpus Adaptation
- 融合卷积、注意力与残差结构,从梅尔频谱中提取关键特征
- 在两大语料库上实现超75%准确率,跨说话人场景显著优于现有方法
- 适合临床评估与康复监测系统开发人员参考使用
失语症语音严重程度分类对运动性语言障碍患者的客观评估与进展监测至关重要。尽管已有方法尝试解决该问题,但在跨说话人(SID)场景下的鲁棒泛化仍具挑战。本文提出DSSCNet,一种结合卷积、压缩-激励(SE)与残差网络的深度神经架构,可从梅尔频谱中提取失语症语音的判别性表征。SE模块聚焦重要特征,减少信息损失并提升性能。同时提出基于检测的迁移学习框架,用于跨语料微调。DSSCNet在两个基准语料库TORGO与UA-Speech上,采用单说话人每严重程度(OSPS)与留一说话人(LOSO)评估协议进行测试。在未微调情况下,分别取得56.84%与62.62%(OSPS)、63.47%与64.18%(LOSO)准确率;经微调后,准确率提升至75.80%与68.25%(OSPS),77.76%与79.44%(LOSO),显著优于现有最优方法,验证了其在多源数据上的有效性与泛化能力。
原文摘要 · Abstract (English)
Dysarthric speech severity classification is crucial for objective clinical assessment and progress monitoring in individuals with motor speech disorders. Although prior methods have addressed this task, achieving robust generalization in speaker-independent (SID) scenarios remains challenging. This work introduces DSSCNet, a novel deep neural architecture that combines Convolutional, Squeeze-Excitation (SE), and Residual network, helping it extract discriminative representations of dysarthric speech from mel spectrograms. The addition of SE block selectively focuses on the important features of the dysarthric speech, thereby minimizing loss and enhancing overall model performance. We also propose a cross-corpus fine-tuning framework for severity classification, adapted from detection-based transfer learning approaches. DSSCNet is evaluated on two benchmark dysarthric speech corpora: TORGO and UA-Speech under speaker-independent evaluation protocols: One-Speaker-Per-Severity (OSPS) and Leave-One-Speaker-Out (LOSO) protocols. DSSCNet achieves accuracies of 56.84% and 62.62% under OSPS and 63.47% and 64.18% under LOSO setting on TORGO and UA-Speech respectively outperforming existing state-of-the-art methods. Upon fine-tuning, the performance improves substantially, with DSSCNet achieving up to 75.80% accuracy on TORGO and 68.25% on UA-Speech in OSPS, and up to 77.76% and 79.44%, respectively, in LOSO. These results demonstrate the effectiveness and generalizability of DSSCNet for fine-grained severity classification across diverse dysarthric speech datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。