用WordNet自动标注阿英词典词义的词性,提升多语言知识库兼容性。
Automatic Part-of-Speech Tagging of Arabic-English Dictionary Senses through WordNet
- 通过英文等价词的词性标签转移实现阿英词义词性标注
- 在阿语词典上达到高准确率,仅依赖少量资源
- 适合缺乏语料和专家资源的语言处理工具开发
本文提出一种算法,用于对阿英双语词典中词义进行词性(POS)标注。该方法应用于Al-Mawrid阿英词典,通过消除歧义后将英文等价词(TEs)的词性标签迁移至词典词义。英文词性来自普林斯顿WordNet。词性标注是将双语词典与WordNet关联或转换为WordNet-LMF格式的前提,其中语义簇(synset)而非单词是基本单元。尽管成本低,但准确率较高。构建NLP/HLT工具通常需要语言学专家、大量投入和长时间。统计方法需大规模标注语料,规则方法需包含丰富语言与世界知识的大词典。这促使了面向低资源语言的轻量级方法出现。
原文摘要 · Abstract (English)
This paper proposed an algorithm for part-of-speech (POS) tagging senses of a bilingual dictionary. The algorithm is applied on the Al-Mawrid Arabic-English dictionary. The tagging task is accomplished by transferring the POS tags of the English translation equivalences (TEs) to the dictionary senses after dis-ambiguities process. The English POS tags of senses are acquired from the Princeton WordNet. POS tagging of bilingual dictionary senses is prerequisite to link a bilingual dictionary to WordNet and/or standardizing that dictionary into WordNet-LMF format where the synset (set of synonyms), not word, is the basic brick. The registered accuracy is high though the cost is little. Building NLP/HLT tools needs linguistic experts, large investments, and long time. For statistical approach, we need large annotated corpora and for rule-based approach, we need large lexicon that contains rich linguistic and world knowledge. That motivates the appearance of what are called resource-light approaches to develop natural language processing (NLP) tools for poor-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。