arXiv:2508.17148cs.CLcs.SD2025-08中稿 · IEEE ASRU 2025被引 3

让语音识别模型学会地理定位,提升对方言口音的识别鲁棒性。

Geolocation-Aware Robust Spoken Language Identification

  • 在自监督模型中加入地理预测任务,用位置信息引导特征学习
  • 在FLEURS上达97.7%准确率,ML-SUPERB方言集提升9.7%
  • 适合需要跨地域语音识别的场景,如多语种语音助手

尽管自监督学习(SSL)显著提升了语音语言识别(LID)性能,现有模型仍难以将同一种语言的不同方言和口音统一归类。为此,我们提出地理定位感知的LID方法,将语言级地理信息融入基于SSL的LID模型。具体地,引入地理预测作为辅助任务,并将预测向量注入中间表示作为条件信号。这种显式条件化促使模型学习更统一的方言与口音表征。在六个多语言数据集上的实验表明,该方法有效提升对语言内变异和未见领域的鲁棒性,在FLEURS上达到97.7%的新最优准确率,在ML-SUPERB 2.0方言集上实现9.7%的相对提升。

原文摘要 · Abstract (English)

While Self-supervised Learning (SSL) has significantly improved Spoken Language Identification (LID), existing models often struggle to consistently classify dialects and accents of the same language as a unified class. To address this challenge, we propose geolocation-aware LID, a novel approach that incorporates language-level geolocation information into the SSL-based LID model. Specifically, we introduce geolocation prediction as an auxiliary task and inject the predicted vectors into intermediate representations as conditioning signals. This explicit conditioning encourages the model to learn more unified representations for dialectal and accented variations. Experiments across six multilingual datasets demonstrate that our approach improves robustness to intra-language variations and unseen domains, achieving new state-of-the-art accuracy on FLEURS (97.7%) and 9.7% relative improvement on ML-SUPERB 2.0 dialect set.

语音识别自监督学习方言识别地理感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。