arXiv:2608.26992cs.LG2026-08

构建首个针对口音英语的快速领域自适应基准,评估预训练语音表征在新口音上的迁移能力。

Benchmarking_Fast_Domain_Adaptation_for_Unsupervised_Speech_Units

论文配图:Benchmarking_Fast_Domain_Adaptation_for_Unsupervised_Speech_Units
图 1 · 摘自论文原文
  • 提出基于多口音数据集的领域自适应评测框架
  • 自适应归一化使跨说话人识别准确率平均提升23.6%
  • 适合研究语音模型鲁棒性与低资源口音适配的研究者

表示学习在下游任务预训练和无监督语音建模中表现优异,但对域外语音(尤其是非标准口音)的适应机制仍不明确。本文提出ABX-Accent基准,基于AESRC数据集包含10种英语口音,每种口音均提供少于10小时的无标签训练数据,并为各口音定制零资源挑战赛的ABX评估指标。以对比预测编码(CPC)模型为基础,采用自适应域归一化进行微调,在男性/女性划分的LibriSpeech上验证方法后,应用于新基准,平均跨说话人ABX得分相对未适配模型提升23.6%。数据与评估指标将在论文接受后开源。

原文摘要 · Abstract (English)

Representation learning has attracted great atten- tion and managed to reach good performances as a pretraining method for downstream tasks or as a first step towards unsu- pervised speech modeling. Yet, little is known about how such methods deal with out-of-domain speech and how could they be adapted in a few shot to new domains. This is important especially for accented speech where one observes a long tail of accents that diverge from the standard ones. We introduce ABX- Accent, a benchmark based on the AESRC dataset that features 10 different accents of English. It includes a small (< 10 hours) unlabelled training set in each of the accents and adaptations of the Zero Resources Challenge ABX evaluation metrics to each of the accents. We illustrate this benchmark with a baseline model that uses adaptive domain normalization to fine tune a pretrained Contrastive Predictive Coding model on the accents. This method is first developed on LibriSpeech using a male/female split. When applied to the new benchmark, the proposed method yields a relative improvement of 23.6% on across-speaker ABX scores on average compared to non adapted models. The data and metrics will be open sourced upon paper acceptance

语音表征领域自适应口音识别无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。