用多模态一致性筛选无参考语音数据,提升方言语音识别效果。
Multimodal Consistency-Guided Reference-Free Data Selection for ASR Accent Adaptation
- 通过扰动解码生成多组伪转录,结合声纹-文本对齐与预测错误率评分
- 在同域设置下选1.5k条数据即达10.91% WER,接近30k标注数据的10.45%
- 适合无标注方言语音适配场景,尤其对抗强口音偏移有效
自动语音识别(ASR)系统在方言语音上表现下降,因声学-音素和语调变化导致与训练数据不匹配,而有标签适配成本高昂。现有伪标签筛选方法多依赖文本信息(如困惑度过滤),易选择流畅但声学不匹配的假设,导致微调时错误放大。为此,本文提出一种基于多模态一致性的无参考数据筛选流程,适用于归纳式、无标签协议下的方言适配。流程首先基于子模互信息进行目标感知预筛选,提升查询相关性并减少下游计算量;随后对每段语音采用扰动解码生成多个伪转录,并利用共享嵌入空间中的声纹-文本对齐及预测词错误率(WER)两个无参考信号进行打分;最后采用简单百分位筛选规则保留可靠伪标签,剔除噪声样本。在同域设置中,从3万条候选中选取约1.5千条,达到10.91% WER,接近使用3万条人工标注的10.45%水平。跨域设置下,经一致性过滤的子集避免了未过滤伪标签在强口音偏移下的性能退化;更强模型的匹配小时实验也验证了其优于随机采样和近期基线方法的增益。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) systems often degrade on accented speech because acoustic-phonetic and prosodic shifts induce a mismatch to training data, making labeled accent adaptation costly. However, common pseudo-label selection heuristics are largely text-centric (e.g., perplexity (PPL) filtering) and can prefer fluent yet acoustically mismatched hypotheses, leading to error amplification when fine-tuning. To address this, we introduce a multimodal consistency-guided, reference-free data selection pipeline for ASR accent adaptation under a transductive, label-free protocol. The pipeline starts with a target-aware preselection step based on submodular mutual information to improve query relevance and reduce downstream computation. It then generates multiple pseudo-transcriptions per utterance via perturbation-based decoding and scores each hypothesis using two reference-free signals: speech--text alignment in a shared embedding space and predicted word error rate (WER). A simple percentile-based selection rule retains reliable pseudo-labels for fine-tuning while discarding noisy utterances. In an in-domain setting, selecting ~1.5k utterances from a 30k pool achieves 10.91% WER, close to 10.45% obtained using 30k supervised labels. In a cross-domain setting with a mismatched candidate pool, consistency-filtered subsets avoid the degradation caused by unfiltered pseudo-labels under strong accent shift, and matched-hour experiments on a stronger ASR backbone further confirm gains over random sampling and recent selection baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。