让语音模型无视说话人差异,精准识别声调,提升低资源方言识别效果
SITA: Learning Speaker-Invariant and Tone-Aware Speech Representations for Low-Resource Tonal Languages
- 分阶段多目标训练:对比学习+声调分离损失,兼顾说话人无关与声调敏感
- 在赫蒙语上跨性别词检索准确率显著提升,自动语音识别性能接近教师模型
- 方法可迁移至普通话,适合作为多语言语音模型的通用声调适配插件
声调类低资源语言使用广泛但语音技术覆盖不足。核心挑战在于学习对性别等干扰因素鲁棒,同时保持对词汇意义关键声调的敏感性。为此,我们提出SITA——一种轻量级适配方案,针对预训练wav2vec风格编码器强化说话人无关性与声调感知能力。SITA采用分阶段多目标训练:(i) 跨性别对比目标促进不同说话人间的词汇一致性,声调排斥损失通过显式分离同词异调实例防止声调坍塌;(ii) 基于连接时序分类(CTC)的辅助语音识别目标结合知识蒸馏,稳定识别相关结构。主要在高度声调化且严重低资源的赫蒙语上评估,现有通用多语言编码器无法有效表示声调。在自建赫蒙词语料库上,SITA显著提升跨性别词检索准确率,同时维持与经语音识别优化的XLS-R教师模型相当的可用语音识别性能。进一步在普通话上观察到类似增益,表明SITA是适用于声调语言的通用、即插即用型适配方法。
原文摘要 · Abstract (English)
Tonal low-resource languages are widely spoken yet remain underserved by modern speech technology. A key challenge is learning representations that are robust to nuisance variation such as gender while remaining tone-aware for different lexical meanings. To address this, we propose SITA, a lightweight adaptation recipe that enforces Speaker-Invariance and Tone-Awareness for pretrained wav2vec-style encoders. SITA uses staged multi-objective training: (i) a cross-gender contrastive objective encourages lexical consistency across speakers, while a tone-repulsive loss prevents tone collapse by explicitly separating same-word different-tone realizations; and (ii) an auxiliary Connectionist Temporal Classification (CTC)-based ASR objective with distillation stabilizes recognition-relevant structure. We evaluate primarily on Hmong, a highly tonal and severely under-resourced language where off-the-shelf multilingual encoders fail to represent tone effectively. On a curated Hmong word corpus, SITA improves cross-gender lexical retrieval accuracy, while maintaining usable ASR accuracy relative to an ASR-adapted XLS-R teacher. We further observe similar gains when transferring the same recipe to Mandarin, suggesting SITA is a general, plug-in approach for adapting multilingual speech encoders to tonal languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。