用词典定义匹配构建更易用的粗粒度词义库。
Coarse-Grained Sense Inventories Based on Semantic Matching between English Dictionaries
- 通过语义匹配融合剑桥词典与WordNet的词义定义。
- 新词义库在语义一致性上优于已有粗粒度词义库。
- 支持CEFR分级,易于扩展,不依赖大规模数据。
WordNet 是最大规模的手工构建概念词典之一,通过语义关系可视化词语间的联系,广泛用于自然语言处理中的词义库存。然而,其细粒度词义常被认为限制了实用性。本文通过语义匹配剑桥词典与WordNet的词义定义,构建新的粗粒度词义库存。我们通过与粗粒度词义库存(Coarse Sense Inventory)的语义一致性对比,验证了所提库存的有效性。新库存的优势包括:对大规模资源依赖低、更优地聚合密切相关词义、支持CEFR等级标注,以及便于扩展与改进。
原文摘要 · Abstract (English)
WordNet is one of the largest handcrafted concept dictionaries visualizing word connections through semantic relationships. It is widely used as a word sense inventory in natural language processing tasks. However, WordNet's fine-grained senses have been criticized for limiting its usability. In this paper, we semantically match sense definitions from Cambridge dictionaries and WordNet and develop new coarse-grained sense inventories. We verify the effectiveness of our inventories by comparing their semantic coherences with that of Coarse Sense Inventory. The advantages of the proposed inventories include their low dependency on large-scale resources, better aggregation of closely related senses, CEFR-level assignments, and ease of expansion and improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。