分析濒危语言达拉瓦尔与107种高资源语言的语音相似性,助力语音技术适配。
Characterization of Speech Similarity Between Australian Aboriginal and High-Resource Languages: A Case Study on Dharawal
- 用预训练多语言语音编码器分析达拉瓦尔与高资源语言的语音相似性。
- 发现达拉瓦尔与拉丁语、毛利语、韩语等有显著语音相似性。
- 为濒危语言语音技术迁移提供数据支持,适合语言保护研究者参考。
澳大利亚原住民语言具有重要的文化和语言价值,但在现代语音人工智能系统中仍严重缺乏代表性。尽管先进的语音基础模型和自动语音识别在高资源环境下表现优异,但往往难以泛化到低资源语言,尤其是缺乏清洁标注语音数据的语言。本文通过精心收集和处理公开录音,构建并清理了达拉瓦尔语(一种低资源澳大利亚原住民语言)的语音数据集。利用该数据集,我们采用预训练的多语言语音编码器,分析达拉瓦尔语与107种高资源语言之间的语音相似性。方法结合(1)误分类率分析以评估语言混淆程度,以及(2)嵌入空间中的余弦相似度和弗雷切特入学距离(FID)进行细粒度相似性测量。实验结果表明,达拉瓦尔语与拉丁语、毛利语、韩语、泰语及威尔士语存在较强的语音相似性。这些发现为未来的迁移学习和模型适应提供了实用指导,并强调了数据收集与基于嵌入的分析在支持濒危语言社区语音技术方面的重要性。
原文摘要 · Abstract (English)
Australian Aboriginal languages are of significant cultural and linguistic value but remain severely underrepresented in modern speech AI systems. While state-of-the-art speech foundation models and automatic speech recognition excel in high-resource settings, they often struggle to generalize to low-resource languages, especially those lacking clean, annotated speech data. In this work, we collect and clean a speech dataset for Dharawal, a low-resource Australian Aboriginal language, by carefully sourcing and processing publicly available recordings. Using this dataset, we analyze the speech similarity between Dharawal and 107 high-resource languages using a pre-trained multilingual speech encoder. Our approach combines (1) misclassification rate analysis to assess language confusability, and (2) fine-grained similarity measurements using cosine similarity and Fréchet Inception Distance (FID) in the embedding space. Experimental results reveal that Dharawal shares strong speech similarity with languages such as Latin, Māori, Korean, Thai, and Welsh. These findings offer practical guidance for future transfer learning and model adaptation efforts, and underscore the importance of data collection and embedding-based analysis in supporting speech technologies for endangered language communities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。