为弱势语言构建结构化词汇数据集的新方法,助力知识公平
L-ReLF: A Framework for Lexical Dataset Creation
- 设计可复现的流程,整合OCR与纠错机制生成高质量词汇数据
- 成果适配Wikidata Lexemes,支持机器翻译等下游任务
- 适合关注低资源语言数字化的学者与开源社区
本文提出L-ReLF(低资源词汇框架),一种可复现的方法论,用于为欠发达语言创建高质量、结构化的词汇数据集。以摩洛哥方言为例,缺乏标准化术语严重阻碍了维基百科等平台的知识平等,常导致编辑依赖不一致的临时方法创建新词。本研究系统性地解决了低资源数据面临的问题,包括来源识别、在光学字符识别(OCR)偏向现代标准阿拉伯语的情况下仍有效利用,以及严格的后处理以纠正错误并统一数据模型。最终生成的结构化数据集完全兼容Wikidata Lexemes,构成关键的技术资源。该方法具有通用性,为其他语言社群提供了构建基础词汇数据的清晰路径,支持机器翻译、形态分析等自然语言处理应用。
原文摘要 · Abstract (English)
This paper introduces the L-ReLF (Low-Resource Lexical Framework), a novel, reproducible methodology for creating high-quality, structured lexical datasets for underserved languages. The lack of standardized terminology, exemplified by Moroccan Darija, poses a critical barrier to knowledge equity in platforms like Wikipedia, often forcing editors to rely on inconsistent, ad-hoc methods to create new words in their language. Our research details the technical pipeline developed to overcome these challenges. We systematically address the difficulties of working with low-resource data, including source identification, utilizing Optical Character Recognition (OCR) despite its bias towards Modern Standard Arabic, and rigorous post-processing to correct errors and standardize the data model. The resulting structured dataset is fully compatible with Wikidata Lexemes, serving as a vital technical resource. The L-ReLF methodology is designed for generalizability, offering other language communities a clear path to build foundational lexical data for downstream NLP applications, such as Machine Translation and morphological analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。