构建大规模多语言形态学数据集,支持实证研究与模型测试。
GLeMM: A large-scale multilingual dataset for morphological research
- 基于维基词典自动化构建,覆盖7种欧洲语言
- 每条词目均标注形态特征,部分含语义描述
- 适合形态学、计算语言学及跨语言研究者使用
在派生形态学中,词形与词义关系的变化机制是什么?现有研究多依赖直觉和小规模数据,难以复现与推广。为此,我们提出GLeMM,一个用于形态学实验与数据驱动描述的大型多语言资源。该数据集具有五大特点:(一)规模大;(二)覆盖德语、英语、西班牙语、法语、意大利语、波兰语、俄语共七种欧洲语言;(三)全自动化构建,跨语言设计一致;(四)每条条目自动标注形态特征;(五)对部分条目编码语义描述。该资源使研究者能够探讨词形与词义在构词中的作用,并开发与验证计算方法。本文介绍GLeMM如何利用维基词典条目构建,并展示多个应用案例。
原文摘要 · Abstract (English)
In derivational morphology, what mechanisms govern the variation in form-meaning relations between words? The answers to this type of questions are typically based on intuition and on observations drawn from limited data, even when a wide range of languages is considered. Many of these studies are difficult to replicate and generalize. To address this issue, we present GLeMM, a new derivational resource designed for experimentation and data-driven description in morphology. GLeMM is characterized by (i) its large size, (ii) its extensive coverage (currently amounting to seven European languages, i.e., German, English, Spanish, French, Italian, Polish, Russian, (iii) its fully automated design, identical across all languages, (iv) the automatic annotation of morphological features on each entry, as well as (v) the encoding of semantic descriptions for a significant subset of these entries. It enables researchers to address difficult questions, such as the role of form and meaning in word-formation, and to develop and experimentally test computational methods that identify the structures of derivational morphology. The article describes how GLeMM is created using Wiktionary articles and presents various case studies illustrating possible applications of the resource.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。