用超网络提升渐冻症患者发音障碍识别准确率
Recognition of Dysarthria in Amyotrophic Lateral Sclerosis patients using Hypernetworks
- 用超网络生成目标网络权重,实现输入条件计算
- 在VOC-ALS数据集上达82.66%准确率,优于基线方法
- 模型参数少、泛化强,适合医疗语音分析场景
肌萎缩侧索硬化症(ALS)是一种进行性神经退行性疾病,常伴随语言清晰度下降。现有研究通过预测临床标准ALSFRS-R来识别发音障碍,依赖特征提取与定制卷积神经网络加全连接层的设计。然而,近期研究表明,采用输入条件计算逻辑的神经网络具有训练更快、性能更好、灵活性高等优势。为此,本文首次将超网络引入发音障碍识别任务。具体地,将音频转换为对数梅尔频谱图、一阶差分与二阶差分,并输入改进的预训练AlexNet模型;随后使用超网络生成目标网络的权重。实验在新收集的公开数据集VOC-ALS上进行,结果表明,该方法准确率达82.66%,超越多种强基线模型,包括多模态融合方法;消融实验验证了该方法的有效性。总体而言,所提方法在泛化能力、参数效率和鲁棒性方面均优于当前最优结果。
原文摘要 · Abstract (English)
Amyotrophic Lateral Sclerosis (ALS) constitutes a progressive neurodegenerative disease with varying symptoms, including decline in speech intelligibility. Existing studies, which recognize dysarthria in ALS patients by predicting the clinical standard ALSFRS-R, rely on feature extraction strategies and the design of customized convolutional neural networks followed by dense layers. However, recent studies have shown that neural networks adopting the logic of input-conditional computations enjoy a series of benefits, including faster training, better performance, and flexibility. To resolve these issues, we present the first study incorporating hypernetworks for recognizing dysarthria. Specifically, we use audio files, convert them into log-Mel spectrogram, delta, and delta-delta, and pass the resulting image through a pretrained modified AlexNet model. Finally, we use a hypernetwork, which generates weights for a target network. Experiments are conducted on a newly collected publicly available dataset, namely VOC-ALS. Results showed that the proposed approach reaches Accuracy up to 82.66% outperforming strong baselines, including multimodal fusion methods, while findings from an ablation study demonstrated the effectiveness of the introduced methodology. Overall, our approach incorporating hypernetworks obtains valuable advantages over state-of-the-art results in terms of generalization ability, parameter efficiency, and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。