arXiv:2506.06785cs.CL2025-06

为1500多种语言的语料添加句法依存标注,助力跨语言类型学研究。

Extending dependencies to the taggedPBC: Word order in transitive clauses

  • 基于已有的词性标注语料,统一添加依存关系标注。
  • 依存结构与三大语言类型数据库中的语序规律高度一致。
  • 适合语言类型学、跨语言比较研究者使用。

taggedPBC(Ring 2025a)包含来自1500多种语言、133个语系和111个孤立语言的超过1800条词性标注的平行语料,覆盖范围远超以往资源,且词性标注准确度良好,可支持预测性跨语言分析(Ring 2025b)。然而该数据集最初未标注依存关系。本文报告了其以CoNLLU格式输出的依存标注版本,将依存信息与词性标注一并扩展至所有语言。尽管标注质量存在争议,但由此得出的及物句中论元与谓词的位置分布,与三大类型学数据库(WALS、Grambank、Autotyp)中专家判定的语序模式高度相关。这表明基于语料的类型学方法(如Baylor et al. 2023;Bjerva 2024)在拓展离散语言类别比较方面具有价值,即使在噪声数据中亦可获得重要洞见,前提是具备足够标注。依存标注语料已通过GitHub公开,供研究与协作使用。

原文摘要 · Abstract (English)

The taggedPBC (Ring 2025a) contains more than 1,800 sentences of pos-tagged parallel text data from over 1,500 languages, representing 133 language families and 111 isolates. While this dwarfs previously available resources, and the POS tags achieve decent accuracy, allowing for predictive crosslinguistic insights (Ring 2025b), the dataset was not initially annotated for dependencies. This paper reports on a CoNLLU-formatted version of the dataset which transfers dependency information along with POS tags to all languages in the taggedPBC. Although there are various concerns regarding the quality of the tags and the dependencies, word order information derived from this dataset regarding the position of arguments and predicates in transitive clauses correlates with expert determinations of word order in three typological databases (WALS, Grambank, Autotyp). This highlights the usefulness of corpus-based typological approaches (as per Baylor et al. 2023; Bjerva 2024) for extending comparisons of discrete linguistic categories, and suggests that important insights can be gained even from noisy data, given sufficient annotation. The dependency-annotated corpora are also made available for research and collaboration via GitHub.

跨语言依存标注类型学语料库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。