用大模型统一提升口吃语音识别与事件检测效果
Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging Recognition and Event Detection
- 大模型联动语音识别与口吃事件检测,动态交互增强理解
- 口吃语音识别错误率降37.71%,事件检测准确率提升46.58%
- 适合语音康复、无障碍技术研究者参考
口吃语音的自动语音识别(ASR)性能受限,影响其在言语康复等领域的应用。本文提出一种基于大语言模型(LLM)的ASR-SED多任务学习框架,联合优化语音识别与口吃事件检测(SED)任务。设计动态交互机制:ASR分支利用CTC生成的软提示辅助大模型上下文建模,SED分支输出口吃嵌入以增强大模型对口吃语音的理解。引入对比学习强化口吃声学特征的区分能力,并采用焦点损失缓解口吃事件类别分布不均问题。在AS-70普通话口吃数据集上的实验表明,该框架将ASR字符错误率(CER)降至5.45%(相对降低37.71%),平均SED F1分数达73.63%(相对提升46.58%)。
原文摘要 · Abstract (English)
The performance bottleneck of Automatic Speech Recognition (ASR) in stuttering speech scenarios has limited its applicability in domains such as speech rehabilitation. This paper proposed an LLM-driven ASR-SED multi-task learning framework that jointly optimized the ASR and Stuttering Event Detection (SED) tasks. We proposed a dynamic interaction mechanism where the ASR branch leveraged CTC-generated soft prompts to assist LLM context modeling, while the SED branch output stutter embeddings to enhance LLM comprehension of stuttered speech. We incorporated contrastive learning to strengthen the discriminative power of stuttering acoustic features and applied Focal Loss to mitigate the long-tailed distribution in stuttering event categories. Evaluations on the AS-70 Mandarin stuttering dataset demonstrated that our framework reduced the ASR character error rate (CER) to 5.45% (-37.71% relative reduction) and achieved an average SED F1-score of 73.63% (+46.58% relative improvement).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。