arXiv:2609.04223cs.CL2026-09

不同语料库导致依赖距离估计差异显著,语言排序可能被扭曲。

How Much Does Corpus Choice Change Dependency-Distance Estimates?

论文配图:How Much Does Corpus Choice Change Dependency-Distance Estimates?
图 1 · 摘自论文原文
  • 用12种预处理方案对比38组同语言树库的依赖距离
  • 近40%的语言排序因语料库更换而反转,语料贡献29%方差
  • 依赖长度最小化规律仍存,但跨语言排名不稳健

从单一语料库得出的依赖距离估计常被视为语言固有属性,但该假设尚未在独立构建的语料库间得到验证。我们利用通用依存标注库v2.18中的38组同语言树库对,通过一致性相关分析、Bland-Altman检验及十二种预处理方案的多宇宙设计,比较了平均依赖距离估计值。跨树库一致性仅处于中等水平:替换一个树库几乎使40%的成对语言排序发生反转,树库选择解释了约29%的组间方差。这一分歧远超单个树库内的抽样误差,并在所有预处理方案下持续存在。然而,每个树库均证实了依赖长度最小化(归一化比值低于1)。数据更支持最小依赖距离(MDD)是语法、语域和标注因素共同作用下的语料依赖型复合特征,而非稳定的语言层级参数:定性上依赖长度最小化普遍成立,但跨语言的序数排名不具备稳定性。

原文摘要 · Abstract (English)

Dependency-distance estimates derived from a single corpus are routinely treated as properties of a language, yet this assumption has not been tested across independently compiled corpora. We compared mean dependency-distance estimates across 38 same-language treebank pairs from Universal Dependencies v2.18, using concordance correlation, Bland-Altman analysis, and a twelve-specification multiverse design. Cross-treebank agreement was moderate at best: substituting one treebank for another reversed nearly 40 percent of pairwise language orderings, and treebank choice accounted for roughly 29 percent of between-group variance. This disagreement substantially exceeded within-treebank sampling error and persisted across all twelve preprocessing specifications. Nevertheless, every treebank confirmed dependency-length minimization (normalized ratio below 1). The data are more consistent with MDD as a corpus-conditioned composite of grammatical, register, and annotation factors than as a stable language-level parameter: the qualitative DLM universal survives corpus substitution, but the ordinal cross-linguistic ranking does not.

依存句法语料库研究语言共性统计偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。