arXiv:2606.19940eess.AS2026-06

对比语言与地理联合监督,发现能更好区分方言地域差异。

Analyzing Language and Geographical Variation in Speech Representations Across 60 Indic Languages

论文配图:Analyzing Language and Geographical Variation in Speech Representations Across 60 Indic Languages
图 1 · 摘自论文原文
  • 用语言+地区联合标注(386类)微调模型,比仅语言标注更优
  • 在嵌入空间中实现语言聚类内有序的地区子结构,地理区分度提升
  • 适合研究印地语系语音表征地理差异的学者参考

自监督语音编码器通常在语言监督下微调,可能忽略地理差异。为分析语言-地区联合监督与仅语言监督下的表征差异,我们对Whisper-base和Wav2Vec2.0-base进行分类任务微调,采用联合语言-地区(386类)和仅语言(60语言)标签。结果表明,语言-地区监督在保持强语言分类能力的同时,显著提升了语言内的地区判别能力。通过归一化条件互信息(NCMI)分析嵌入结构,发现联合监督使全球语言簇内部形成有序的地区子簇,增强地理可分性,且不破坏语言层级组织。

原文摘要 · Abstract (English)

Self-supervised speech encoders are often fine-tuned with language supervision, which can overlook geographical variation. To understand the learned representations under joint supervision of language and district compared to language-only supervision, we fine-tune Whisper-base and Wav2Vec2.0-base for classification tasks with joint language-district (386 classes) and language-only classification (60 languages). The language-district supervision improves district discrimination conditioned on language in the embedding space while strong marginal language classification. We analyze the structure of the learned embeddings using Normalized Conditional Mutual Information (NCMI), showing that language-district supervision produces global language clusters with structured within language subclusters aligned to district variation, enhancing geographical separability without degrading language-level organization.

语音表征地理差异多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。