用语音难度分指导采样,让语音识别更懂口吃等非标准发音。
Data-Efficient ASR Personalization for Non-Normative Speech Using an Uncertainty-Based Phoneme Difficulty Score for Guided Sampling
- 通过变分低秩适配估算语音模型不确定性,生成音素难度评分。
- 在英德语数据集上,准确率提升显著,且困难音素稳定持续。
- 适合语音识别、康复医学及残障人士语音交互研究者。
语音识别系统在非标准发音场景下表现不佳,主要因声学差异大且数据稀缺。本文提出一种数据高效的个性化方法,利用音素级不确定性引导微调。不依赖计算成本高的集成模型,而是采用变分低秩适配(VI LoRA)估算基础模型的主观不确定性,形成综合音素难度评分(PhDScore),进而驱动针对性过采样策略。在英语和德语数据集上评估,包括对两年内两次临床报告的纵向分析,结果表明:(1) 基于 VI LoRA 的不确定性与专家临床评估更一致,优于标准熵;(2) PhDScore 能捕捉稳定持久的发音困难;(3) 不确定性引导采样显著提升受损语音的识别准确率。
原文摘要 · Abstract (English)
ASR systems struggle with non-normative speech due to high acoustic variability and data scarcity. We propose a data-efficient method using phoneme-level uncertainty to guide fine-tuning for personalization. Instead of computationally expensive ensembles, we leverage Variational Low-Rank Adaptation (VI LoRA) to estimate epistemic uncertainty in foundation models. These estimates form a composite Phoneme Difficulty Score (PhDScore) that drives a targeted oversampling strategy. Evaluated on English and German datasets, including a longitudinal analysis against two clinical reports taken one year apart, we demonstrate that: (1) VI LoRA-based uncertainty aligns better with expert clinical assessments than standard entropy; (2) PhDScore captures stable, persistent articulatory difficulties; and (3) uncertainty-guided sampling significantly improves ASR accuracy for impaired speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。