arXiv:2509.00058cs.AI2025-09

对比四类语音障碍检测模型,揭示准确率、可控性与可解释性的权衡。

A Comparative Study of Controllability, Explainability, and Performance in Dysfluency Detection Models

  • 对比 YOLO-Stutter、FluentNet、UDM、SSDM 四种模型的性能与可解释性。
  • UDM 在准确率和临床可解释性上表现最佳,适合实际应用。
  • 实验发现 SSDM 无法复现,提示模型可靠性需谨慎评估。

近期语音障碍检测研究提出了多种建模范式,从轻量级目标检测模型(YOLOStutter)到模块化可解释框架(UDM)。尽管基准数据集上的性能持续提升,但临床应用不仅需要高准确率,还需模型具备可控性和可解释性。本文系统比较了四种代表性方法——YOLO-Stutter、FluentNet、UDM 和 SSDM——在性能、可控性与可解释性三个维度的表现。通过多数据集评估及临床专家打分,发现 YOLO-Stutter 和 FluentNet 具有高效简洁优势,但透明度有限;UDM 在准确率与临床可解释性间取得最佳平衡;而 SSDM 虽具潜力,但在本实验中未能完全复现。分析揭示了各方法间的权衡关系,并指明未来临床可用模型的发展方向。同时提供各方法的实现细节与部署建议。

原文摘要 · Abstract (English)

Recent advances in dysfluency detection have introduced a variety of modeling paradigms, ranging from lightweight object-detection inspired networks (YOLOStutter) to modular interpretable frameworks (UDM). While performance on benchmark datasets continues to improve, clinical adoption requires more than accuracy: models must be controllable and explainable. In this paper, we present a systematic comparative analysis of four representative approaches--YOLO-Stutter, FluentNet, UDM, and SSDM--along three dimensions: performance, controllability, and explainability. Through comprehensive evaluation on multiple datasets and expert clinician assessment, we find that YOLO-Stutter and FluentNet provide efficiency and simplicity, but with limited transparency; UDM achieves the best balance of accuracy and clinical interpretability; and SSDM, while promising, could not be fully reproduced in our experiments. Our analysis highlights the trade-offs among competing approaches and identifies future directions for clinically viable dysfluency modeling. We also provide detailed implementation insights and practical deployment considerations for each approach.

语音障碍可解释性模型对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。