构建超大规模多语言语料库,助力跨语言语言学研究
The taggedPBC: Annotating a massive parallel corpus for crosslinguistic investigations
- 构建包含1940+语言的标注平行语料库,覆盖155个语系
- 标签准确率与顶尖工具及人工标注数据高度一致
- 提出新指标N1比值,可有效预测未知语言的词序类型
现有跨语言研究数据集通常要么涵盖少量语言的海量数据,要么覆盖大量语言的有限数据,限制了对人类语言普遍性的揭示。尽管一些项目已尝试构建多语言标注语料库,但仍受资源制约。本文介绍的taggedPBC是迄今最大的标注平行语料库,涵盖超过1,940种语言,涉及155个语言家族和78个孤立语言,远超以往资源。该数据集中特定标签的准确性与高资源语言的SOTA标注器(SpaCy、Trankit)以及人工标注语料库(Universal Dependencies Treebanks)表现高度一致。此外,基于该数据集提出的N1比率与三大类型学数据库(WALS、Grambank、AUTOYP)中专家判定的不及物句词序具有显著相关性,使用该特征训练的高斯朴素贝叶斯分类器能准确识别未收录语言的基本不及物句词序。尽管仍需进一步扩展,taggedPBC为基于语料的跨语言研究提供了重要基础,已通过GitHub开源供学术合作使用。
原文摘要 · Abstract (English)
Existing datasets available for crosslinguistic investigations have tended to focus on large amounts of data for a small group of languages or a small amount of data for a large number of languages. This means that claims based on these datasets are limited in what they reveal about universal properties of the human language faculty. While this has begun to change through the efforts of projects seeking to develop tagged corpora for a large number of languages, such efforts are still constrained by limits on resources. The current paper reports on a large tagged parallel dataset which has been developed to partially address this issue. The taggedPBC contains POS-tagged parallel text data from more than 1,940 languages, representing 155 language families and 78 isolates, dwarfing previously available resources. The accuracy of particular tags in this dataset is shown to correlate well with both existing SOTA taggers for high-resource languages (SpaCy, Trankit) as well as hand-tagged corpora (Universal Dependencies Treebanks). Additionally, a novel measure derived from this dataset, the N1 ratio, correlates with expert determinations of intransitive word order in three typological databases (WALS, Grambank, AUTOYP) such that a Gaussian Naive Bayes classifier trained on this feature can accurately identify basic intransitive word order for languages not in those databases. While much work is still needed to expand and develop this dataset, the taggedPBC is an important step to enable corpus-based crosslinguistic investigations, and is made available for research and collaboration via GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。