arXiv:2510.27241cs.CL2025-10ACL被引 2

发现语言信息存在周期性,且超出句子等常规结构边界。

Identifying the Periodicity of Information in Natural Language

  • 用突变率序列检测文本信息周期性,识别出隐藏的规律模式。
  • 大量人类语言呈现显著周期性,部分周期超越句子等传统单位。
  • 方法可助力大模型生成内容检测,适合研究语言结构与生成模型者。

自然语言信息密度的最新理论提出了一个关键问题:自然语言在编码信息中有多强的周期性?本文提出一种名为自适应突变率周期检测(AutoPeriod of Surprisal, APS)的新方法,采用标准周期性检测算法,可识别单篇文档中突变率序列的显著周期。对多个语料库的应用结果显示:首先,相当比例的人类语言展现出强烈的信息周期性;其次,发现了超出典型文本结构单元(如句法边界、基本话语单元等)分布范围的新周期,并通过谐波回归建模进一步验证。结论表明,语言信息的周期性是结构因素与长距离驱动因素共同作用的结果。本文还讨论了该方法的优势及其在大语言模型生成内容检测中的潜力。

原文摘要 · Abstract (English)

Recent theoretical advancement of information density in natural language has brought the following question on desk: To what degree does natural language exhibit periodicity pattern in its encoded information? We address this question by introducing a new method called AutoPeriod of Surprisal (APS). APS adopts a canonical periodicity detection algorithm and is able to identify any significant periods that exist in the surprisal sequence of a single document. By applying the algorithm to a set of corpora, we have obtained the following interesting results: Firstly, a considerable proportion of human language demonstrates a strong pattern of periodicity in information; Secondly, new periods that are outside the distributions of typical structural units in text (e.g., sentence boundaries, elementary discourse units, etc.) are found and further confirmed via harmonic regression modeling. We conclude that the periodicity of information in language is a joint outcome from both structured factors and other driving factors that take effect at longer distances. The advantages of our periodicity detection method and its potentials in LLM-generation detection are further discussed.

语言周期性信息密度大模型检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。