arXiv:2603.18612cs.CLcs.SD2026-03

构建多语言无监督音素发现基准,验证语音离散单元与音素的关联性。

DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units

  • 用10小时语音数据,从离散语音单元中无监督发现音素
  • 多语言测试显示模型生成单元与音素相关性良好,但跨语言差异明显
  • 提供4个预训练模型,适合语音分析与音系学研究者使用

我们提出DiscoPhon,一个用于评估从离散语音单元中无监督发现音素的多语言基准。该基准涵盖6种开发语言和6种测试语言,覆盖广泛的音位对比。在仅提供10小时未见过的语言语音数据条件下,系统需将离散单元映射到预定义的音素库,通过多对一或一对一分配实现。结果序列在单元质量、识别和分割方面进行评估。我们提供了四个预训练的多语言HuBERT和SpidR基线模型,并表明当前模型中的离散语音单元已包含足够的音素信息,能与音素产生良好相关性,但不同语言间存在差异。

原文摘要 · Abstract (English)

We introduce DiscoPhon, a multilingual benchmark for evaluating unsupervised phoneme discovery from discrete speech units. DiscoPhon covers 6 dev and 6 test languages, chosen to span a wide range of phonemic contrasts. Given only 10 hours of speech in a previously unseen language, systems must produce discrete units that are mapped to a predefined phoneme inventory, through either a many-to-one or a one-to-one assignment. The resulting sequences are evaluated for unit quality, recognition and segmentation. We provide four pretrained multilingual HuBERT and SpidR baselines, and show that phonemic information is available enough in current models for derived units to correlate well with phonemes, though with variations across languages.

语音分析音素发现无监督学习多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。