arXiv:2606.24324cs.CL2026-06中稿 · LREC 2026被引 3

构建了涵盖语义与篇章关系的捷克语语料库,支持多领域自然语言处理研究。

Prague Dependency Treebank -- Consolidated 2.0: Enriching a Complex Annotation Scheme

论文配图:Prague Dependency Treebank -- Consolidated 2.0: Enriching a Complex Annotation Scheme
图 1 · 摘自论文原文
  • 统一标注捷克语语料,整合语义与篇章关系等多层信息
  • 建成近400万词元、覆盖多种文体的高质量语料库
  • 适合语言学研究者及需多层级标注数据的NLP开发者

普拉格依存语料库框架独特地系统性整合了语言的不同层面,包括带有多种句间现象(特别是指代和话语关系)的意义表示。本文介绍其第二版统一版本(PDT-C 2.0),该项目历时近30年,最终形成一个统一、连贯标注、涵盖多种语体、约400万词元的捷克语语料库,并配套完全兼容的词典。该丰富标注语料库不仅支撑持续的语言学研究,也广泛用于传统与新型自然语言处理工具的国际比较,以及向其他形式化体系转换。语料库与训练好的解析器均以CC BY-NC-SA许可开放。

原文摘要 · Abstract (English)

The Prague Dependency Treebank framework is unique in its attempt to systematically include and link different layers of language, including a meaning representation with several types of inter-sentential phenomena, especially coreference and discourse relations. We present its second consolidated version (PDT-C 2.0), which concludes almost 30-years long project of sustained development of the resource to a uniformly and coherently annotated, genre-diversified, almost 4 million token language resource of Czech language, with accompanying fully compatible lexicons. In addition to continuous linguistic research, the richly linguistically annotated corpus is also widely used in international comparisons of the development of traditional and novel NLP tools as well as in conversions into other formalisms. The corpus and the trained parsers are available under the CC BY-NC-SA licence.

语料库依存标注语言学捷克语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。