用40分钟短语音数据,让濒危语言实现可用的语音识别。
How I Built ASR for Endangered Languages with a Spoken Dictionary
- 用简短发音词典替代长语句标注,降低数据要求。
- 仅需40分钟数据,曼岛盖尔语识别错误率低于50%。
- 适合无大量标注数据的濒危语言复兴项目使用。
全球近半语言处于濒危状态。自动语音识别(ASR)是语言复兴的核心技术,但多数语言因缺乏语句级监督数据而无法支持。例如,曼岛盖尔语(约2200名使用者)虽自1948年起就有转录语音,却仍不被现代系统支持。本文探索构建濒危语言ASR所需的最小数据量及形式。结果表明,短格式发音资源即可作为有效替代方案:仅需40分钟此类数据,便可在曼岛盖尔语上实现可用的ASR(WER < 50%)。我们进一步将该方法应用于另一濒危语言康沃尔语(约600名使用者),验证其可复现性。研究显示,进入门槛在数据量与形式上远低于以往认知,为无法承担苛刻数据要求的语言社区带来希望。
原文摘要 · Abstract (English)
Nearly half of the world's languages are endangered. Speech technologies such as Automatic Speech Recognition (ASR) are central to revival efforts, yet most languages remain unsupported because standard pipelines expect utterance-level supervised data. Speech data often exist for endangered languages but rarely match these formats. Manx Gaelic ($\sim$2,200 speakers), for example, has had transcribed speech since 1948, yet remains unsupported by modern systems. In this paper, we explore how little data, and in what form, is needed to build ASR for critically endangered languages. We show that a short-form pronunciation resource is a viable alternative, and that 40 minutes of such data produces usable ASR for Manx ($<$50\% WER). We replicate our approach, applying it to Cornish ($\sim$600 speakers), another critically endangered language. Results show that the barrier to entry, in quantity and form, is far lower than previously thought, giving hope to endangered language communities that cannot afford to meet the requirements arbitrarily imposed upon them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。