arXiv:2606.11531cs.CLcs.IT2026-06

用分层复用模式衡量语言复杂度,发现不同语言复杂度相近。

Measuring language complexity from hierarchical reuse of recurring patterns

论文配图:Measuring language complexity from hierarchical reuse of recurring patterns
图 1 · 摘自论文原文
  • 基于算法信息论,通过分层复用重复结构计算语言复杂度。
  • 21个语种数据中复杂度基本不变,远低于语料长度变化幅度。
  • 结果支持语言复杂度守恒与认知加工机制的关联性。

我们提出梯路指数(ladderpath index),一种基于算法信息论的语言复杂度度量方法。该指数通过计算重构序列所需的最少步骤,量化分层复用重复子结构的能力,是一种可精确计算但受限的算法压缩形式,与柯尔莫哥洛夫复杂度相关但不同。我们在来自并行通用依存库的21个平行语料上应用此方法,发现梯路指数在不同语言间近似不变,且变化幅度远小于语料长度差异。当所有语料映射到统一二进制表示后,这一现象更明显,为语言等复杂度假说提供了独立于表示方式的证据。我们还观察到字符集大小与语料长度之间、词汇级与语料级重构复杂度之间的权衡关系,支持总复杂度守恒并在语言层级间重新分配的权衡假说。梯路方法识别出的可复用子结构无需语言先验,与自然词汇中的词和词素重合。其分层复用机制与认知科学提出的认知组块化相似,表明人类语言处理共享记忆与计算约束下的嵌套可复用单位生成机制。这一联系为等复杂度与权衡假说提供了新解释,将其根植于跨语言共通的认知架构。

原文摘要 · Abstract (English)

We introduce the ladderpath index as a measure of language complexity grounded in algorithmic information theory. It counts the minimum steps needed to reconstruct a sequence through hierarchical reuse of repeated substructures, capturing an exactly computable but constrained form of algorithmic compressibility related to, but distinct from, Kolmogorov complexity. We apply the ladderpath approach to 21 parallel corpora from the Parallel Universal Dependencies dataset. The ladderpath index is approximately invariant across the languages, and varies much less than the corpus length. This is more pronounced when all corpora are mapped to a unified binary representation, providing evidence for the equi-complexity hypothesis from a representation-independent perspective. We also observe trade-offs between character inventory size and corpus length, and between vocabulary-level and corpus-level reconstruction complexity, supporting the trade-off hypothesis that total complexity is conserved and redistributed across linguistic levels. The reusable substructures identified by the ladderpath approach, without any linguistic input, overlap with words and morphological components attested in the natural vocabulary. The hierarchical reuse captured by the ladderpath approach parallels the chunking mechanisms proposed in cognitive science, where the human cognitive system compresses linguistic input into nested, reusable units under shared memory and processing constraints. This connection between cognitive chunking and the ladderpath approach provides a new interpretation for the equi-complexity and trade-off hypotheses, grounding both in the shared cognitive architecture that underlies language processing across human languages.

语言复杂度算法信息论认知科学分层结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。