零样本检测语音不流畅,同时转写音素并识别异常
Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection
- 用WFST框架在不训练下同步转写音素与检测不流畅
- 在模拟和真实数据上达到最优的音素错误率和检测率
- 轻量可解释,适合临床语音病理诊断场景
自动检测语音不流畅有助于言语语言病理学家高效转录障碍性语音,提升诊断与治疗规划。传统方法多限于分类,提供临床洞察不足,且文本无关模型在上下文依赖情况下误判严重。本文提出Dysfluent-WFST,一种无需训练的零样本解码器,可同时转写音素并检测不流畅。该框架兼容WavLM等上游编码器,无需额外训练。在模拟与真实语音数据上均达到音素错误率与不流畅检测的最先进性能。研究证明,解码中显式建模发音行为比复杂架构更关键,该方法轻量、可解释、有效。
原文摘要 · Abstract (English)
Automatic detection of speech dysfluency aids speech-language pathologists in efficient transcription of disordered speech, enhancing diagnostics and treatment planning. Traditional methods, often limited to classification, provide insufficient clinical insight, and text-independent models misclassify dysfluency, especially in context-dependent cases. This work introduces Dysfluent-WFST, a zero-shot decoder that simultaneously transcribes phonemes and detects dysfluency. Unlike previous models, Dysfluent-WFST operates with upstream encoders like WavLM and requires no additional training. It achieves state-of-the-art performance in both phonetic error rate and dysfluency detection on simulated and real speech data. Our approach is lightweight, interpretable, and effective, demonstrating that explicit modeling of pronunciation behavior in decoding, rather than complex architectures, is key to improving dysfluency processing systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。