arXiv:2509.14304cs.SDcs.AI2025-09

UDM模型兼顾高精度与可解释性,助力临床语音治疗落地。

Deploying UDM Series in Real-Life Stuttered Speech Applications: A Clinical Evaluation Framework

  • 模块化设计+音素对齐,输出可解释的失语检测结果。
  • F1达0.89,临床可解释性评分4.2/5.0,诊断时间减少34%。
  • 适合需高可信度AI辅助的言语治疗临床场景。

失语和口吃语音检测系统长期面临准确率与临床可解释性之间的权衡。尽管端到端深度学习模型性能优异,但其黑箱特性限制了临床应用。本文聚焦伯克利开发的当前最优框架——无约束失语建模(UDM)系列,该框架融合模块化结构、显式音素对齐与可解释输出,适用于真实临床部署。通过涉及患者和认证言语语言病理学家(SLPs)的大量实验,我们证实UDM在保持顶尖性能(F1: 0.89±0.04)的同时,提供4.2/5.0的临床可解释性评分。部署研究显示,87%的临床医生接受该系统,诊断时间减少34%。结果表明,UDM为临床环境中人工智能辅助语音治疗提供了切实可行的路径。

原文摘要 · Abstract (English)

Stuttered and dysfluent speech detection systems have traditionally suffered from the trade-off between accuracy and clinical interpretability. While end-to-end deep learning models achieve high performance, their black-box nature limits clinical adoption. This paper looks at the Unconstrained Dysfluency Modeling (UDM) series-the current state-of-the-art framework developed by Berkeley that combines modular architecture, explicit phoneme alignment, and interpretable outputs for real-world clinical deployment. Through extensive experiments involving patients and certified speech-language pathologists (SLPs), we demonstrate that UDM achieves state-of-the-art performance (F1: 0.89+-0.04) while providing clinically meaningful interpretability scores (4.2/5.0). Our deployment study shows 87% clinician acceptance rate and 34% reduction in diagnostic time. The results provide strong evidence that UDM represents a practical pathway toward AI-assisted speech therapy in clinical environments.

语音识别临床应用可解释性失语检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。