arXiv:2508.16681cs.AIcs.CL2025-08

提出可解释的语音迟滞检测框架,提升临床可用性。

Revisiting Rule-Based Stuttering Detection: A Comprehensive Analysis of Interpretable Models for Clinical Applications

  • 基于多级声学特征与分级决策结构,增强规则系统性能。
  • 延音检测准确率达97%-99%,在不同语速下表现稳定。
  • 适合需要可审计、可调参的临床场景,兼容现代AI流程。

口吃影响全球约1%的人群,损害沟通与生活质量。尽管深度学习推动了自动言语不流畅检测的进展,但在临床应用中,可解释性与透明性仍使基于规则的方法不可或缺。本文综合分析了多个语料库(UCLASS、FluencyBank、SEP-28k)中的规则型口吃检测系统,提出一种改进的规则框架,包含语速归一化、多层级声学特征分析和分层决策结构。该方法在保持完全可解释性的前提下达到有竞争力的性能,尤其在延音检测上准确率达97%-99%,且在不同语速下表现稳定。此外,我们证明这些可解释模型可作为现代机器学习流程中的候选生成器或约束模块,弥合传统语音病理学与当代AI系统的差距。分析表明,尽管神经方法在无约束环境下可能略高精度,但规则方法在需决策可审计、患者个性化调整与实时反馈的临床场景中具有独特优势。

原文摘要 · Abstract (English)

Stuttering affects approximately 1% of the global population, impacting communication and quality of life. While recent advances in deep learning have pushed the boundaries of automatic speech dysfluency detection, rule-based approaches remain crucial for clinical applications where interpretability and transparency are paramount. This paper presents a comprehensive analysis of rule-based stuttering detection systems, synthesizing insights from multiple corpora including UCLASS, FluencyBank, and SEP-28k. We propose an enhanced rule-based framework that incorporates speaking-rate normalization, multi-level acoustic feature analysis, and hierarchical decision structures. Our approach achieves competitive performance while maintaining complete interpretability-critical for clinical adoption. We demonstrate that rule-based systems excel particularly in prolongation detection (97-99% accuracy) and provide stable performance across varying speaking rates. Furthermore, we show how these interpretable models can be integrated with modern machine learning pipelines as proposal generators or constraint modules, bridging the gap between traditional speech pathology practices and contemporary AI systems. Our analysis reveals that while neural approaches may achieve marginally higher accuracy in unconstrained settings, rule-based methods offer unique advantages in clinical contexts where decision auditability, patient-specific tuning, and real-time feedback are essential.

语音识别可解释性临床应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。