机器评估口吃严重程度为何不如医生?原因在感知与统计的鸿沟。
Bridging the Perceptual-Statistical Gap in Dysarthria Assessment: Why Machine Learning Still Falls Short
- 提出'感知-统计鸿沟'概念,解释模型与专家差异根源
- 发现标注噪声和医生评分不一致是性能瓶颈
- 建议用语音感知特征、自监督预训练等方法改进
自动语音分析在口吃检测与严重程度评估中具有重要临床潜力。尽管声学建模和深度学习进展迅速,模型性能仍不及人类专家。本文系统分析该差距成因,提出‘感知-统计鸿沟’这一核心概念。通过剖析人类专家的听觉感知过程,回顾机器学习表征与建模策略,梳理现有特征集与方法,并理论分析标签噪声与评分者间变异对性能的限制。进一步提出实践方案:设计感知驱动特征、自监督预训练、以语音识别为指导的目标函数、多模态融合、人机协同训练及可解释性方法。最后,建议采用符合临床目标的实验协议与评估指标,推动未来研究向可信赖、可解释的口吃评估工具发展。
原文摘要 · Abstract (English)
Automated dysarthria detection and severity assessment from speech have attracted significant research attention due to their potential clinical impact. Despite rapid progress in acoustic modeling and deep learning, models still fall short of human expert performance. This manuscript provides a comprehensive analysis of the reasons behind this gap, emphasizing a conceptual divergence we term the ``perceptual-statistical gap''. We detail human expert perceptual processes, survey machine learning representations and methods, review existing literature on feature sets and modeling strategies, and present a theoretical analysis of limits imposed by label noise and inter-rater variability. We further outline practical strategies to narrow the gap, perceptually motivated features, self-supervised pretraining, ASR-informed objectives, multimodal fusion, human-in-the-loop training, and explainability methods. Finally, we propose experimental protocols and evaluation metrics aligned with clinical goals to guide future research toward clinically reliable and interpretable dysarthria assessment tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。