多语言框架同时解决失语症检测、分级、语音转写与清晰语音生成。
A Multilingual Framework for Dysarthria: Detection, Severity Classification, Speech-to-Text, and Clean Speech Generation
- 统一AI框架处理六项任务,跨语言训练提升泛化能力。
- 检测与分级准确率均达97%,模型注意力聚焦低谐波区域。
- 跨语言迁移有效降低英文低资源场景的重建误差。
失语症是一种导致言语缓慢且难以理解的运动性言语障碍,严重影响交流。该病常见于帕金森病和肌萎缩侧索硬化症等神经疾病,但现有工具缺乏跨语言和严重程度的通用性。本文提出一个统一的多语言人工智能框架,涵盖六项关键任务:(1) 二分类失语症检测,(2) 严重程度分类,(3) 清晰语音生成,(4) 语音转文字,(5) 情绪识别,(6) 语音克隆。分析了英语、俄语和德语数据集,通过频谱可视化与声学特征提取指导模型训练。二分类检测模型在三语言上均达97%准确率,证明强跨语言泛化能力。严重程度分类模型测试准确率达97%,可解释性显示模型注意力集中于低谐波成分。基于俄语配对数据训练的语音重建管道,训练损失0.03,测试损失0.06;在英语数据上微调后,训练损失降至0.02,测试损失0.03,凸显跨语言迁移学习在低资源场景的潜力。语音转写管道经三轮训练后词错误率(WER)为0.1367,实现对失语语音的精准转录,支持后续情绪识别与语音克隆。整体成果可助力多语言失语症诊断与患者沟通改善。
原文摘要 · Abstract (English)
Dysarthria is a motor speech disorder that results in slow and often incomprehensible speech. Speech intelligibility significantly impacts communication, leading to barriers in social interactions. Dysarthria is often a characteristic of neurological diseases including Parkinson's and ALS, yet current tools lack generalizability across languages and levels of severity. In this study, we present a unified AI-based multilingual framework that addresses six key components: (1) binary dysarthria detection, (2) severity classification, (3) clean speech generation, (4) speech-to-text conversion, (5) emotion detection, and (6) voice cloning. We analyze datasets in English, Russian, and German, using spectrogram-based visualizations and acoustic feature extraction to inform model training. Our binary detection model achieved 97% accuracy across all three languages, demonstrating strong generalization across languages. The severity classification model also reached 97% test accuracy, with interpretable results showing model attention focused on lower harmonics. Our translation pipeline, trained on paired Russian dysarthric and clean speech, reconstructed intelligible outputs with low training (0.03) and test (0.06) L1 losses. Given the limited availability of English dysarthric-clean pairs, we fine-tuned the Russian model on English data and achieved improved losses of 0.02 (train) and 0.03 (test), highlighting the promise of cross-lingual transfer learning for low-resource settings. Our speech-to-text pipeline achieved a Word Error Rate of 0.1367 after three epochs, indicating accurate transcription on dysarthric speech and enabling downstream emotion recognition and voice cloning from transcribed speech. Overall, the results and products of this study can be used to diagnose dysarthria and improve communication and understanding for patients across different languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。