通过OCR技术自动化提取印尼地方语言文档,降低资源构建成本。
DriveThru: a Document Extraction Platform and Benchmark Datasets for Indonesian Local Language Archives
- 利用OCR技术自动提取印刷文献中的印尼地方语言内容。
- 相比传统方法,字符准确率和词准确率显著提升。
- 适合需要低成本构建小语种NLP资源的研究者使用。
印度尼西亚是语言多样性最丰富的国家之一,但其本土语言在自然语言处理(NLP)研究与技术中仍严重缺失。尽管过去两年已有若干针对印尼语的NLP资源建设尝试,但多数依赖人工标注,难以规模化扩展至更多语言。许多印尼语言虽无网络存在,却在书籍、杂志、报纸等纸质文献中保存良好。本论文提出一种新方法:通过数字化这些现有印刷资源,构建语言数据集。我们开发了DriveThru平台,采用光学字符识别(OCR)技术实现文档内容自动提取,大幅减少人工投入与成本。同时,本研究评估了当前最先进的大语言模型(LLM)在后置OCR纠错中的表现,结果表明其可有效提升字符准确率(CAR)和词准确率(WAR),优于通用OCR工具。
原文摘要 · Abstract (English)
Indonesia is one of the most diverse countries linguistically. However, despite this linguistic diversity, Indonesian languages remain underrepresented in Natural Language Processing (NLP) research and technologies. In the past two years, several efforts have been conducted to construct NLP resources for Indonesian languages. However, most of these efforts have been focused on creating manual resources thus difficult to scale to more languages. Although many Indonesian languages do not have a web presence, locally there are resources that document these languages well in printed forms such as books, magazines, and newspapers. Digitizing these existing resources will enable scaling of Indonesian language resource construction to many more languages. In this paper, we propose an alternative method of creating datasets by digitizing documents, which have not previously been used to build digital language resources in Indonesia. DriveThru is a platform for extracting document content utilizing Optical Character Recognition (OCR) techniques in its system to provide language resource building with less manual effort and cost. This paper also studies the utility of current state-of-the-art LLM for post-OCR correction to show the capability of increasing the character accuracy rate (CAR) and word accuracy rate (WAR) compared to off-the-shelf OCR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。