arXiv:2506.01775cs.CL2025-06中稿 · Comput-EL 2025 Wor…被引 2

用混合方法让百年原住民文字图像可机器阅读

Developing a Mixed-Methods Pipeline for Community-Oriented Digitization of Kwak'wala Legacy Texts

  • 结合现成OCR与语言识别,分离并识别古老文本
  • 实现11卷扫描文本的高精度转录,准确率显著提升
  • 为濒危语言数字化提供可复用的技术路径

Kwak'wala是加拿大不列颠哥伦比亚省的一种原住民语言,拥有超过百年的文献记录,现存活跃的使用者、教师和学习者群体。已有11卷由弗朗茨·博厄斯与乔治·亨特合作创作的早期文献被扫描,但因手写体复杂而无法被机器读取。本文采用最新的光学字符识别(OCR)技术,对仅以图像形式存在的Kwak'wala文本进行处理,探讨了实际应用中的挑战及必要调整。在前人基础上,提出融合现成OCR、语言识别与掩码技术,有效分离出Kwak'wala文本,并通过后处理校正模型生成高质量转录结果,为该语言的现代正字法转写及其他语言技术开发奠定基础。

原文摘要 · Abstract (English)

Kwak'wala is an Indigenous language spoken in British Columbia, with a rich legacy of published documentation spanning more than a century, and an active community of speakers, teachers, and learners engaged in language revitalization. Over 11 volumes of the earliest texts created during the collaboration between Franz Boas and George Hunt have been scanned but remain unreadable by machines. Complete digitization through optical character recognition has the potential to facilitate transliteration into modern orthographies and the creation of other language technologies. In this paper, we apply the latest OCR techniques to a series of Kwak'wala texts only accessible as images, and discuss the challenges and unique adaptations necessary to make such technologies work for these real-world texts. Building on previous methods, we propose using a mix of off-the-shelf OCR methods, language identification, and masking to effectively isolate Kwak'wala text, along with post-correction models, to produce a final high-quality transcription.

语言复兴OCR原住民语言文本数字化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。