arXiv:2505.21093eess.AScs.SD2025-05被引 2

用音视频+机器学习客观评估渐冻症患者言语障碍,准确率达93%

Multimodal Assessment of Speech Impairment in ALS Using Audio-Visual and Machine Learning Approaches

  • 融合音频视频特征,用极端梯度提升模型预测言语障碍程度
  • 在5-25分量表上实现0.93的均方根误差,优于单一模态
  • 适合临床早期筛查和居家监测,降低评估成本

渐冻症患者的言语分析是评估延髓功能障碍的有效工具,但现有临床方法多依赖主观判断或昂贵设备。本研究结合音视频分析与机器学习,利用少量语音任务的声学与运动学特征,训练并测试多种回归模型。最佳结果来自采用多模态特征的极端梯度提升回归器,在5至25分量表上达到0.93的均方根误差。结果表明,音视频融合分析可显著提升言语障碍评估的客观性,为早期检测与家庭环境下的延髓功能监测提供可行工具。

原文摘要 · Abstract (English)

The analysis of speech in individuals with amyotrophic lateral sclerosis is a powerful tool to support clinicians in the assessment of bulbar dysfunction. However, current methods used in clinical practice consist of subjective evaluations or expensive instrumentation. This study investigates different approaches combining audio-visual analysis and machine learning to predict the speech impairment evaluation performed by clinicians. Using a small dataset of acoustic and kinematic features extracted from audio and video recordings of speech tasks, we trained and tested some regression models. The best performance was achieved using the extreme boosting machine regressor with multimodal features, which resulted in a root mean squared error of 0.93 on a scale ranging from 5 to 25. Results suggest that integrating audio-video analysis enhances speech impairment assessment, providing an objective tool for early detection and monitoring of bulbar dysfunction, also in home settings.

渐冻症语音评估多模态机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。