用大模型生成更自然的口吃语音数据,提升检测准确率
Analysis and Evaluation of Synthetic Data Generation in Speech Dysfluency Detection
- 用大语言模型增强口吃模拟,生成11类口吃语音
- 在自建数据集上实现当前最佳检测效果
- 适合语音病理分析与自动诊断研究者使用
口吃检测对临床诊断和语言评估至关重要,但现有方法受限于高质量标注数据稀缺。尽管语音合成技术进步使口吃语音生成成为可能,但现有合成数据普遍存在语调不自然、上下文多样性不足的问题。为此,我们提出 LLM-Dys——首个基于大语言模型增强的口吃语音综合语料库,涵盖11类跨词与音素层面的口吃类型。依托该资源,我们改进了端到端口吃检测框架,实验验证其达到当前最优性能。所有数据、模型与代码已开源至 https://github.com/Berkeley-Speech-Group/LLM-Dys。
原文摘要 · Abstract (English)
Speech dysfluency detection is crucial for clinical diagnosis and language assessment, but existing methods are limited by the scarcity of high-quality annotated data. Although recent advances in TTS model have enabled synthetic dysfluency generation, existing synthetic datasets suffer from unnatural prosody and limited contextual diversity. To address these limitations, we propose LLM-Dys -- the most comprehensive dysfluent speech corpus with LLM-enhanced dysfluency simulation. This dataset captures 11 dysfluency categories spanning both word and phoneme levels. Building upon this resource, we improve an end-to-end dysfluency detection framework. Experimental validation demonstrates state-of-the-art performance. All data, models, and code are open-sourced at https://github.com/Berkeley-Speech-Group/LLM-Dys.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。