用可靠性加学习性加权,提升太平洋原住民语音识别效果
QuaSR: Quality-Aware Sample Reweighting for Pacific Indigenous Speech Recognition
- 结合声学、字幕和对齐可靠性评估数据质量,用训练损失衡量模型学习难易度
- 在4种太平洋原住民语言上验证,加权后识别准确率显著提升
- 适合低资源语音识别场景,尤其对标注不一致的数据集有明显改进
为低资源语言训练自动语音识别(ASR)模型面临数据稀缺与标注质量参差的问题。特别是太平洋原住民语音语料库常存在异质声学条件、转录不一致以及声文对齐可靠性差异,导致标准微调方法对噪声或误导性监督信号敏感。本文提出QuaSR,一种简单而有效的样本重加权框架,融合数据侧可靠性与模型侧可学习性以改善ASR适应性能。具体地,从声学、转录和对齐三个维度估计数据可靠性,同时利用模型训练损失衡量学习难易度。两者结合生成统一的样本效用分数,用于指导训练权重。我们在四种太平洋原住民语言上进行评估,结果表明该效用分数与模型适应性能高度相关。此外,QuaSR在所有语言上均优于标准微调和替代数据选择策略,展现出通过难度评分提升低资源语音学习的新路径。
原文摘要 · Abstract (English)
Training automatic speech recognition (ASR) models for low-resource languages is challenging due to limited data and highly variable supervision quality. In particular, Pacific Indigenous speech corpora often exhibit heterogeneous acoustic conditions, transcript inconsistencies, and varying degrees of acoustic-text alignment reliability, making standard fine-tuning approaches sensitive to noisy or misleading supervision signals. In this work, we propose QuaSR, a simple yet effective weighting framework that combines data-side reliability with model-side learnability to improve ASR adaptation. Specifically, we estimate data reliability from acoustic, transcription, and alignment, while measuring learnability using training loss from the model. These two complementary signals are integrated into a unified sample utility score to produce training weights for the samples. We also evaluated across four Pacific Indigenous languages, which shows that the proposed utility scores reliably correlate with adaptation performance. Furthermore, QuaSR consistently improves ASR adaptation over standard fine-tuning and alternative data selection strategies, highlighting a new way to leverage difficulty scores for low-resource speech learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。