让语音识别更准确记录口吃和流畅性修饰,提升临床研究价值
On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping Artifacts
- 用轻量级方法将口吃和流畅性特征转为特殊标记,实现精准识别
- 在英德双语口吃数据集上验证有效,德国数据仍存性能差距
- 揭示分词器对英语的偏见,适合语音诊断与无障碍技术研究者
自动语音识别在处理口吃语音时仍面临挑战,即使采用现代端到端(E2E)系统也常忽略口吃和流畅性修饰,导致转录不完整,临床与研究价值有限。本文提出一种参数高效的适配方法,将口吃和流畅性修改作为特殊标记融入转录结果,在模拟(LibriStutter,英语)和真实(KSoF,德语)口吃语音数据集上进行评估。为缓解英语主导带来的性能差异与偏差,引入语言自适应预训练的多步微调策略。分词分析进一步揭示分词器存在英语中心偏见,影响德语数据表现。结果表明,轻量适配技术可有效提升口吃感知的语音识别能力,同时暴露多语言端到端系统的关键局限。
原文摘要 · Abstract (English)
Automatic transcription of stuttered speech remains a challenge, even for modern end-to-end (E2E) automatic speech recognition (ASR) frameworks. Dysfluencies and fluency-shaping artifacts are often overlooked, resulting in non-verbatim transcriptions with limited clinical and research value. We propose a parameter-efficient adaptation method to decode dysfluencies and fluency modifications as special tokens within transcriptions, evaluated on simulated (LibriStutter, English) and natural (KSoF, German) stuttered speech datasets. To mitigate ASR performance disparities and bias towards English, we introduce a multi-step fine-tuning strategy with language-adaptive pretraining. Tokenization analysis further highlights the tokenizer's English-centric bias, which poses challenges for improving performance on German data. Our findings demonstrate the effectiveness of lightweight adaptation techniques for dysfluency-aware ASR while exposing key limitations in multilingual E2E systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。