arXiv:2509.17367cs.CL2025-09

法律文本有独特复杂性,现有大模型无法完全复现。

Scale-free Characteristics of Multilingual Legal Texts and the Limitations of LLMs

  • 用词汇增长、词频波动等指标量化文本复杂度
  • 法律条文词汇增长慢、术语一致性高,生成文本更像普通语言
  • 适合研究法律文本特征或大模型局限的学者

我们通过尺度不变指标对比了不同领域文本的复杂性。采用海普斯指数β(词汇增长)、泰勒指数α(词频波动)和压缩率r(冗余度)、熵来量化语言复杂度。研究涵盖三个领域:法律文件(法规、判例、契约)、通用自然语言文本(文学、维基百科)和AI生成文本(GPT)。结果表明,法律文本的词汇增长较慢(β更低),术语一致性更高(α更高),其中法规的β最低、α最高,反映其严格起草规范;而判例和契约则表现出更高的β和更低的α。相比之下,GPT生成文本的统计特征更接近通用语言模式。这说明法律文本具有特定结构与复杂性,当前生成模型尚无法充分还原。

原文摘要 · Abstract (English)

We present a comparative analysis of text complexity across domains using scale-free metrics. We quantify linguistic complexity via Heaps' exponent $β$ (vocabulary growth), Taylor's exponent $α$ (word-frequency fluctuation scaling), compression rate $r$ (redundancy), and entropy. Our corpora span three domains: legal documents (statutes, cases, deeds) as a specialized domain, general natural language texts (literature, Wikipedia), and AI-generated (GPT) text. We find that legal texts exhibit slower vocabulary growth (lower $β$) and higher term consistency (higher $α$) than general texts. Within legal domain, statutory codes have the lowest $β$ and highest $α$, reflecting strict drafting conventions, while cases and deeds show higher $β$ and lower $α$. In contrast, GPT-generated text shows the statistics more aligning with general language patterns. These results demonstrate that legal texts exhibit domain-specific structures and complexities, which current generative models do not fully replicate.

法律文本语言复杂性大模型局限

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。