研究大模型能否模仿语言的分形复杂性,发现其输出易暴露生成痕迹。
A Tale of Two Structures: Do LLMs Capture the Fractal Complexity of Language?
- 通过信息论分析语言与大模型输出的分形特性差异。
- 自然语言分形参数集中,模型输出则波动大,可作检测依据。
- 结果对不同模型架构均成立,适合文本生成与检测研究者参考。
语言在信息论复杂度(即每标记比特数)上表现出分形结构,具有跨尺度自相似性和长程依赖性。本文研究大语言模型(LLMs)是否能复现此类分形特征,并识别温度设置和提示方法等导致失败的条件。我们发现,自然语言的分形参数集中在狭窄范围内,而模型输出的参数变化广泛,表明分形参数可用于识别非平凡的模型生成文本。值得注意的是,这些发现及其他结论对多种模型架构(如Gemini 1.0 Pro、Mistral-7B、Gemma-2B)均稳健。我们还发布了包含超过24万篇由不同大模型生成的文本数据集(涵盖预训练与指令微调模型,不同解码温度与提示方法),以及对应的真人撰写文本。本工作揭示了分形特性、提示策略与统计模仿之间的复杂互动,为生成、评估和检测合成文本提供了洞见。
原文摘要 · Abstract (English)
Language exhibits a fractal structure in its information-theoretic complexity (i.e. bits per token), with self-similarity across scales and long-range dependence (LRD). In this work, we investigate whether large language models (LLMs) can replicate such fractal characteristics and identify conditions-such as temperature setting and prompting method-under which they may fail. Moreover, we find that the fractal parameters observed in natural language are contained within a narrow range, whereas those of LLMs' output vary widely, suggesting that fractal parameters might prove helpful in detecting a non-trivial portion of LLM-generated texts. Notably, these findings, and many others reported in this work, are robust to the choice of the architecture; e.g. Gemini 1.0 Pro, Mistral-7B and Gemma-2B. We also release a dataset comprising of over 240,000 articles generated by various LLMs (both pretrained and instruction-tuned) with different decoding temperatures and prompting methods, along with their corresponding human-generated texts. We hope that this work highlights the complex interplay between fractal properties, prompting, and statistical mimicry in LLMs, offering insights for generating, evaluating and detecting synthetic texts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。