将捷克语多领域树库转换为通用依存标注格式
Meet UD_Czech-PDTC: A Large and Genre-Rich Treebank in Universal Dependencies

- 将普拉哈依赖树库整合版转换为通用依存标注规范
- 新资源规模超原版两倍,涵盖更广语体与领域
- 适合语言学研究与多语言依存分析任务使用
捷克语自2015年通用依存标注(Universal Dependencies, UD)首版发布以来一直被收录,且长期是代表性最强的语言之一。原始普拉哈依赖树库(PDT)规模远超多数其他UD树库。近年来,三个源自普拉哈家族的新数据集被加入,并经过全面重注释,形成“普拉哈依赖树库-整合版”(PDT-C)。相比原版PDT,PDT-C规模超过两倍,语体与领域多样性显著提升。本文描述该资源向通用依存标注的转换过程。尽管两种标注体系表面相似,但在依存结构拓扑、词性标注粒度及关系类型体系方面存在诸多细微差异。文中通过实例展示这些差异,讨论其背后的设计动机,并提出转换阶段的解决策略。我们认为,虽然PDT在通用性上略逊于标准UD,但其多层级标注信息丰富,足以支持基础UD树构建,并提供更多有用细节。
原文摘要 · Abstract (English)
Czech has been part of Universal Dependencies since its first release in 2015. It has also been one of the best represented languages, with the Prague Dependency Treebank being order of magnitude larger than most other UD treebanks. More recently, three other datasets from the Prague family were added and the annotations thoroughly revisited, forming the "Prague Dependency Treebank-Consolidated" (PDT-C). In comparison to the original PDT, PDT-C is more than twice as large, but it is also much more diverse in terms of genres and domains. In this paper, we describe the conversion of the new resource to Universal Dependencies. While the two annotation schemes are relatively similar at the first sight, there are numerous small differences in topology of the dependency structures and in granularity of the POS and relation type inventories. We demonstrate a selection of such differences on examples, discuss the diverging motivations, as well as ways to overcome the differences during conversion. We argue that while PDT is less "universal" and more tightly bound to one language, its multi-layer annotation is rich and provides all information needed for basic UD trees, and much more.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。