用模式压缩词形字典,让百万级词条高效管理。
MorfFlex: Handling Rich Morphology

- 用规则模式替代重复词条,大幅压缩字典体积
- 支持超过1亿词形、100万词元的高覆盖率
- 适合需要一致标注的语料库和先进NLP工具
我们提出MorfFlex,一种适用于具有丰富屈折与派生规则的语言的词形字典架构。以捷克语为例,MorfFlex CZ 是一个包含超过1亿词形和100万词元的词形字典,以简单无结构的<词形, 词元, 标注>三元组形式分发,但其手动维护的源文件和转换脚本中编码了复杂的屈折与派生模式。这些模式显著降低了字典规模。该字典是普鲁士依存句法树库人工标注一致性的重要保障,并支撑了如MorphoDiTa等前沿自动工具。本文重点介绍如何有效管理字典中的复杂形态系统,并展示该语言资源在保持语料标注一致性及支持高级NLP应用方面的实用性。
原文摘要 · Abstract (English)
We present MorfFlex, a morphological dictionary architecture suitable for languages with extensive regularity in both inflection and derivation. As the primary example of MorfFlex in use we introduce MorfFlex CZ, a morphological dictionary of Czech. It is distributed as a simple, unstructured list of <wordform, lemma, tag> triplets, however, its manually maintained, unpublished source files and conversion scripts encode a sophisticated system of inflectional and derivational patterns. These patterns dramatically reduce the otherwise enormous size of the dictionary, which currently contains over 100 million wordforms and more than 1 million lemmas. The MorfFlex CZ dictionary serves as an essential resource for ensuring the consistency of manual morphological annotation in the Prague Dependency Treebanks and underpins state-of-the-art automatic tools such as MorphoDiTa. In this paper, we focus on: (i) presenting an effective method for managing the rich morphological system within the dictionary, and (ii) demonstrating the utility of such a language resource for maintaining annotation consistency in corpora and supporting the development of advanced NLP applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。