将阿拉伯语-英语词典数字化为可计算资源,提升机器处理能力。
Analyzing and Encoding the Al-Mawrid Arabic-English Dictionary with the ISO Language Markup Framework and TEI Lex-0
- 采用国际标准框架整合双语词典结构与语义信息。
- 解析字母'ع'样本,实现91%结构解析准确率,同义词召回率达98%。
- 适合阿拉伯语NLP与数字人文研究者参考,支持开放数据链接。
本文提出一种系统化方法,将阿语-英語词典《Al-Mawrid》从传统印刷品转化为标准化计算词库,填补阿拉伯语词汇基础设施空白。研究结合国际标准化组织的词汇标记框架(ISO LMF)与文本编码倡议(TEI Lex-0)双重标准,通过编辑视角处理词典的宏观与微观结构,解决20世纪双语词典常见的结构模糊与标点不一致问题。基于对代表性样本(字母‘ع’,占全书4.6%)的实证分析,验证了编码流程的科学性,实现结构解析准确率91%。信息抽取规则经量化评估,同义词提取达85%精确率与98%召回率,其他形态语义特征精确率达88%。论文还对比现有阿拉伯语词汇资源,指出TEI Lex-0在处理隐含‘开放集’语义关系及分散形态线索时的局限性。此外,研究构建基于前缀的引用体系,推动该资源融入语言学关联开放数据(LLOD),实现可扩展的语义网集成。最终成果为阿拉伯语自然语言处理与数字人文领域提供可复现、可互操作的数字化工作流。
原文摘要 · Abstract (English)
This paper presents a robust methodology for the systematic digitization and encoding of the Al-Mawrid Arabic-English dictionary, transforming it from a legacy print resource into a standardized computational lexicon. Addressing a significant gap in Arabic lexical infrastructure, the study adopts a dual-standard framing that aligns the ISO Lexical Markup Framework (LMF) with the Text Encoding Initiative TEI Lex-0 guidelines. By applying an editorial view to the dictionary's macro- and microstructure, the research resolves the structural ambiguities and punctuation inconsistencies typical of 20th-century bilingual dictionaries. The methodology is grounded in an empirical analysis of the dictionary's lexical knowledge density. Drawing on a representative sample (the letter Ayn, comprising 4.6% of the total volume), the study provides scientific weight to the encoding process, demonstrating a structural parsing accuracy of 91%. Quantitative evaluation of the information extraction rules reveals high performance, with 85% precision and 98% recall for synonyms, and 88% precision for other morpho-semantic features. Beyond technical description, the paper provides a critical comparison with existing Arabic lexical resources and discusses the limitations of TEI Lex-0 when modelling specific Arabic phenomena, such as implicit "open set" semantic relations and scattered morphological cues. Furthermore, the study explores the potential for Linguistic Linked Open Data (LLOD) integration by establishing a scalable prefix-based referencing system that facilitates the resource's inclusion in the semantic web. The result is an interoperable, machine-tractable resource that provides a reproducible workflow for the retro-digitization of complex legacy bilingual lexicons within the Arabic NLP and Digital Humanities communities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。