LLM写作缺乏人类随时间演变的风格变化,存在明显的时间扁平化现象。
Temporal Flattening in LLM-Generated Text: Comparing Human and LLM Writing Trajectories
- 构建跨时段人类与LLM写作轨迹对比数据集,涵盖三类文本。
- LLM生成文本词汇多样性更高,但语义与情感漂移显著低于人类。
- 仅凭时间变异模式即可94%准确区分人类与LLM写作,适合检测合成文本。
大型语言模型(LLMs)在日常应用中广泛使用,如内容生成与代码编写,每次交互均视为无状态,独立生成回应。然而人类写作具有内在的纵向特性:作者的写作风格和认知状态随月、年演变。这引出核心问题:LLMs能否在长时间跨度上复现此类时间结构?我们构建并公开发布了一个纵向数据集,包含412位人类作者的6,086篇文档,覆盖2012–2024年三个领域(学术摘要、博客、新闻),并与三种代表性LLM在标准及带历史条件生成设置下的轨迹进行比较。通过语义、词汇和认知-情感表征的漂移与方差度量,发现LLM生成文本存在时间扁平化现象:尽管词汇多样性更高,但语义与认知-情感漂移显著低于人类。这些差异极具预测性:仅基于时间变异性模式,即可实现94%准确率与98% ROC-AUC区分人类与LLM轨迹。结果表明,无论是否引入增量历史,时间扁平化始终存在,揭示了当前部署范式的根本局限。该差距对需真实时间结构的应用(如合成训练数据、纵向文本建模)有直接影响。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in daily applications, from content generation to code writing, where each interaction treats the model as stateless, generating responses independently without memory. Yet human writing is inherently longitudinal: authors' styles and cognitive states evolve across months and years. This raises a central question: can LLMs reproduce such temporal structure across extended time periods? We construct and publicly release a longitudinal dataset of 412 human authors and 6,086 documents spanning 2012--2024 across three domains (academic abstracts, blogs, news) and compare them to trajectories generated by three representative LLMs under standard and history-conditioned generation settings. Using drift and variance-based metrics over semantic, lexical, and cognitive-emotional representations, we find temporal flattening in LLM-generated text. LLMs produce greater lexical diversity but exhibit substantially reduced semantic and cognitive-emotional drift relative to humans. These differences are highly predictive: temporal variability patterns alone achieve 94% accuracy and 98% ROC-AUC in distinguishing human from LLM trajectories. Our results demonstrate that temporal flattening persists regardless of whether LLMs generate independently or with access to incremental history, revealing a fundamental property of current deployment paradigms. This gap has direct implications for applications requiring authentic temporal structure, such as synthetic training data and longitudinal text modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。