arXiv:2607.25284eess.AScs.SD2026-07中稿 · publication at Int…

用语音嵌入构建患者图谱,预测渐冻症严重程度与进展。

Multi-Phonation Graph Learning with Self-Supervised Speech Embeddings for ALS Detection and Progression Prediction

论文配图:Multi-Phonation Graph Learning with Self-Supervised Speech Embeddings for ALS Detection and Progression Prediction
图 1 · 摘自论文原文
  • 将2秒语音片段的自监督嵌入构建成患者级图结构
  • 在SAND数据集上达到0.73(严重度)和0.69(进展)的宏F1
  • 适合低资源环境下渐冻症早期筛查与动态监测

肌萎缩侧索硬化症(ALS)逐步损害言语运动控制,使声学分析成为评估病情严重程度和进展的潜在生物标志物。本文提出一种基于个体的图神经网络框架,将多个发音录音聚合为由预训练自监督语音嵌入(SSL)构建的k近邻图,使用2秒片段。在SAND数据集上对比四种前端模型(wav2vec 2.0、HuBERT、data2vec-audio、UniSpeech-SAT)和五种图神经网络(GCN、残差GCN、GAT、GraphSAGE、GIN),任务包括5类构音障碍严重度分类与4类ALSFRS-R进展预测。在官方验证集上,最佳配置(HuBERT+GIN)分别取得0.73与0.69的宏F1,优于基准方法(0.61与0.58)。结果表明,结合图神经网络与预训练跨语言语音表示在低资源条件下具有显著潜力。

原文摘要 · Abstract (English)

Amyotrophic lateral sclerosis (ALS) progressively impairs speech motor control, making acoustic analysis a promising biomarker for severity and progression estimation. We propose a subject-level graph framework that aggregates multiple phonation recordings into a unique k-nearest-neighbor graph built from pretrained SSL embeddings of 2s segments. We compare four SSL front-ends (wav2vec 2.0, HuBERT, data2vec-audio, and UniSpeech-SAT) and five graph neural networks (GCN, residual GCN, GAT, GraphSAGE, and GIN) on the SAND dataset tasks (339 participants: 205 ALS, 134 control): 5-class dysarthria severity and 4-class ALSFRS-R progression prediction. On the official validation set, the best configuration (HuBERT+GIN) achieves macro-F$_1$ of 0.73 for Task 1 and 0.69 for Task 2, outperforming SAND validation baselines (0.61 and 0.58). These results highlight the potential of combining GNNs with pretrained cross-lingual speech representations for low-resource ALS detection and progression monitoring.

ALS检测语音分析图神经网络自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。