提出注意力选择性合并方法,提升儿童语音识别在低资源下的表现
Selective Attention Merging for low resource tasks: A case study of Child ASR
- 通过选择性合并注意力矩阵中的任务向量,优化模型知识迁移
- 在MyST数据集上相对词错误率降低14%,达到8.69的新纪录
- 适合低资源语音识别研究者,尤其关注儿童语音场景
尽管语音基础模型(SFMs)在多种语音任务中表现优异,但在儿童自动语音识别(ASR)等低资源任务上的性能受限于预训练数据不足。为此,本文探索了多种模型融合技术,以利用在更大、更丰富语料库上训练的模型知识。本文提出一种新方法——选择性注意力(SA)合并,通过有选择地合并注意力矩阵中的任务向量,提升SFMs在低资源任务中的表现。在MyST数据库上的实验显示,该方法可使相对词错误率最多降低14%,优于现有模型融合与数据增强技术。结合数据增强与SA合并,我们在Whisper-small模型上实现了MyST数据库上8.69的最新词错误率,凸显了SA合并提升低资源ASR的潜力。
原文摘要 · Abstract (English)
While Speech Foundation Models (SFMs) excel in various speech tasks, their performance for low-resource tasks such as child Automatic Speech Recognition (ASR) is hampered by limited pretraining data. To address this, we explore different model merging techniques to leverage knowledge from models trained on larger, more diverse speech corpora. This paper also introduces Selective Attention (SA) Merge, a novel method that selectively merges task vectors from attention matrices to enhance SFM performance on low-resource tasks. Experiments on the MyST database show significant reductions in relative word error rate of up to 14%, outperforming existing model merging and data augmentation techniques. By combining data augmentation techniques with SA Merge, we achieve a new state-of-the-art WER of 8.69 on the MyST database for the Whisper-small model, highlighting the potential of SA Merge for improving low-resource ASR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。