两阶段微调提升失语症语音识别个性化效果
Two-Stage Adaptation for Non-Normative Speech Recognition: Revisiting Speaker-Independent Initialization for Personalization
- 先用多说话人非典型语音数据做无关说话人微调,再进行说话人特定微调
- 在AphasiaBank和UA-Speech上均实现更优识别准确率,且对正常语音泛化影响小
- 适合需要个性化适配非典型语音的医疗语音识别场景
为失语症等非典型语音定制自动语音识别系统极具挑战。尽管说话人特定微调(SS-FT)被广泛采用,通常直接从通用预训练模型初始化。然而,在语音不匹配情况下,说话人无关适应是否能提供更强的初始化尚不明确。本文提出两阶段适应框架:首先在多说话人非典型语音数据上进行说话人无关微调(SI-FT),随后进行说话人特定微调(SS-FT)。通过在相同每说话人条件下与直接SS-FT的对照实验,验证了该方法在AphasiaBank和UA-Speech数据集上使用Whisper-Large-v3和Qwen3-ASR模型时,持续提升个性化性能,同时保持可接受的域外(OOD)泛化能力。此外在典型语音数据集TED-LIUM v3和FLEURS上的评估也表明该方法具有良好的鲁棒性。
原文摘要 · Abstract (English)
Personalizing automatic speech recognition (ASR) systems for non-normative speech, such as dysarthric and aphasic speech, is challenging. While speaker-specific fine-tuning (SS-FT) is widely used, it is typically initialized directly from a generic pre-trained model. Whether speaker-independent adaptation provides a stronger initialization prior under such mismatch remains unclear. In this work, we propose a two-stage adaptation framework consisting of speaker-independent fine-tuning (SI-FT) on multi-speaker non-normative data followed by SS-FT, and evaluate it through a controlled comparison with direct SS-FT under identical per-speaker conditions. Experiments on AphasiaBank and UA-Speech with Whisper-Large-v3 and Qwen3-ASR, alongside evaluation on typical-speech datasets TED-LIUM v3 and FLEURS, show that two-stage adaptation consistently improves personalization while maintaining manageable out-of-domain (OOD) trade-offs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。