构建跨语言语音资源,助力儿童语音建模与音系研究
IPA-CHILDES & G2P+: Feature-Rich Resources for Cross-Lingual Phonology and Phonemic Language Modeling
- 开发G2P+工具,基于Phoible数据库实现一致音位表征转换
- 创建31语言儿童语料库IPA CHILDES,覆盖自发性口语与亲子对话
- 验证音位分布可跨语言学习音系特征,适用于多语言语音建模
本文介绍两项资源:(i) G2P+,一个将拼写数据转换为一致音位表示的工具;(ii) IPA CHILDES,涵盖31种语言的儿童中心语音音位数据集。现有图符转音素工具生成的音位词汇表常与既定音位体系不一致,G2P+通过利用Phoible数据库中的音位库存解决此问题。借助该工具,我们对CHILDES进行音位标注,形成IPA CHILDES。该数据集填补了现有音位数据集在多语言覆盖、自发性言语及亲子语言关注方面的空白。我们通过在11种语言上训练音位语言模型并探测其音位特征,发现音位的分布特性足以跨语言学习主要类别和发音部位特征。
原文摘要 · Abstract (English)
In this paper, we introduce two resources: (i) G2P+, a tool for converting orthographic datasets to a consistent phonemic representation; and (ii) IPA CHILDES, a phonemic dataset of child-centered speech across 31 languages. Prior tools for grapheme-to-phoneme conversion result in phonemic vocabularies that are inconsistent with established phonemic inventories, an issue which G2P+ addresses by leveraging the inventories in the Phoible database. Using this tool, we augment CHILDES with phonemic transcriptions to produce IPA CHILDES. This new resource fills several gaps in existing phonemic datasets, which often lack multilingual coverage, spontaneous speech, and a focus on child-directed language. We demonstrate the utility of this dataset for phonological research by training phoneme language models on 11 languages and probing them for distinctive features, finding that the distributional properties of phonemes are sufficient to learn major class and place features cross-lingually.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。