arXiv:2512.02669cs.SDcs.AI2025-12被引 1

对比四种方法,用语音数据分类帕金森病导致的言语障碍严重程度。

SAND Challenge: Four Approaches for Dysartria Severity Classification

  • 用视觉变压器、一维卷积、双向LSTM和集成学习分别建模语音特征。
  • 基于声门与共振峰特征的集成模型在验证集上达到0.86的宏平均F1最优。
  • 深度学习模型虽性能稍逊,但提供可解释性洞察,适合临床辅助诊断。

本文对四种不同的建模方法在语音分析神经退行性疾病(SAND)挑战赛中用于分类言语障碍严重程度进行了统一研究。所有模型均使用相同的语音录音数据集完成五分类任务。我们评估了:(1) 基于频谱图图像的视觉变压器(ViT-OF)方法;(2) 八个一维卷积神经网络结合多数投票融合的1D-CNN方法;(3) 九个双向LSTM模型通过多数投票融合的BiLSTM-OF方法;(4) 两级学习框架下结合声门与共振峰特征的分层XGBoost集成模型。每种方法均有详细描述,并在包含53名说话人的验证集上进行性能比较。结果表明,尽管基于特征工程的XGBoost集成模型取得最高宏平均F1为0.86,深度学习模型(ViT、CNN、BiLSTM)也实现了0.70的可观F1分数,并提供了互补性的问题理解视角。

原文摘要 · Abstract (English)

This paper presents a unified study of four distinct modeling approaches for classifying dysarthria severity in the Speech Analysis for Neurodegenerative Diseases (SAND) challenge. All models tackle the same five class classification task using a common dataset of speech recordings. We investigate: (1) a ViT-OF method leveraging a Vision Transformer on spectrogram images, (2) a 1D-CNN approach using eight 1-D CNN's with majority-vote fusion, (3) a BiLSTM-OF approach using nine BiLSTM models with majority vote fusion, and (4) a Hierarchical XGBoost ensemble that combines glottal and formant features through a two stage learning framework. Each method is described, and their performances on a validation set of 53 speakers are compared. Results show that while the feature-engineered XGBoost ensemble achieves the highest macro-F1 (0.86), the deep learning models (ViT, CNN, BiLSTM) attain competitive F1-scores (0.70) and offer complementary insights into the problem.

言语障碍分类语音分析集成学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。