arXiv:2606.15510cs.CLcs.DL2026-06

首个跨8个时期的开源希腊语依存句法语料库,支持多语言对照分析。

AthDGC: An Open Diachronic Greek Treebank with Indo-European Parallels

  • 基于PROIEL标准,统一标注八时期希腊语文本
  • 实现新约与拉丁文等四语种的逐行对齐
  • 适合历史语言学、多语言对比研究者使用

AthDGC(“雅典-PROIEL”)是一个开放的端到端工作流与语料库。据我们所知,它是首个公开许可的希腊语依存句法语料库,覆盖从古风期、古典期、通用希腊语、晚期古代、拜占庭、晚期拜占庭、早期现代到现代希腊语共八个历时阶段,采用统一的PROIEL XML 2.0模式,并实现新约文本与拉丁文(武尔加塔)、哥特文(沃尔菲拉)、古教会斯拉夫文(马里亚努斯)及古典亚美尼亚文之间的逐节对齐。标注基于经PROIEL训练的Stanford Stanza流程;句子级对齐使用多语言句向量模型LaBSE;词级对齐采用multilingual-BERT注意力机制结合AwesomeAlign方法。v0.4版本提供精选样本与开源工具包;完整标注语料库分片仍处于希腊国家超算中心v0.5审核阶段。量化规模、各文献每节数量及各时期标注行数详见v0.5发布说明,待审核完成后更新。概念DOI:10.5281/zenodo.20439182。

原文摘要 · Abstract (English)

AthDGC ("Athens-PROIEL") is an open, end-to-end workflow and dataset. It is, to the best of our knowledge, the first openly licensed dependency-parsed treebank of Greek that spans eight diachronic periods, namely Archaic, Classical, Koine, Late Antique, Byzantine, Late Byzantine, Early Modern, and Modern Greek, under a single PROIEL XML 2.0 schema, with verse-level cross-alignment of the New Testament to Latin (Vulgate), Gothic (Wulfila), Old Church Slavonic (Marianus), and Classical Armenian. AthDGC builds on the PROIEL Treebank Family (Haug and Johndal 2008; Eckhoff et al. 2018), which established the schema and the Koine-Greek reference set for the project. Annotation uses the Stanford Stanza PROIEL-trained workflow; sentence-level alignment uses LaBSE, a multilingual sentence-embedding model; word-level alignment uses multilingual-BERT attention through the AwesomeAlign procedure. The v0.4 release provides curated samples and the open-source toolkit; the full annotated corpus partitions remain under v0.5 audit on the Greek national HPC. Quantitative scale, per-witness verse counts, and per-period annotated-row counts are reported in the v0.5 release notes, after the audit pass completes. Concept DOI: 10.5281/zenodo.20439182.

历史语言学依存句法多语言对齐开源语料

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。