arXiv:2607.17228cs.CL2026-07

发现大模型文本的词组分布模式,揭示其风格与语义的内在关联。

Literary Non-Style in LLM-Generated Text

  • 分析大模型生成文本中词组的统计分布规律。
  • 发现高阶词组与语义内容高度相关,风格缺陷源于语义受限。
  • 适合关注大模型文本质量与人类写作风格差异的研究者。

以往研究显示,大模型生成文本在数量和质量上均与人类写作存在差异。大模型文本在风格上具有独特特征,导致其文本呈现特定的“感觉”,且其语义范围远窄于人类。本文发现大模型生成文本中存在简单而稳定的词组(n-grams)统计分布模式。通过质性分析这些词组,揭示了大模型在风格上的不足。由于高阶词组与语义内容密切相关,因此得出结论:风格与语义问题并非可清晰分离。

原文摘要 · Abstract (English)

Prior work on LLM-generated text has demonstrated quantitative and qualitative departures from text produced by humans. LLM-generated texts differ from human writing in style, resulting in a characteristic textual "feel," while the semantic range of LLMs is much restricted compared to that of humans. In this contribution, I note simple but consistent patterns in the statistical distribution of n-grams within LLM-generated text. Via qualitative analysis of these n-grams, I reveal deficiencies in LLM style. Because higher-order n-grams correlate to semantic content, I conclude that questions of style and semantics are not cleanly separable.

大模型生成文本风格语义分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。