arXiv:2508.16385cs.CL2025-08被引 2

AI写作有独特语言指纹,可被识别为非人类产物

ChatGPT-generated texts show authorship traits that identify them as non-human

  • 用风格分析和多维语域检测对比人类与AI文本
  • 模型在不同场景下风格变化有限,词汇偏好更倾向名词
  • 语法结构差异揭示人类思维特征,适合安全与内容审核应用

大型语言模型能模仿多种写作风格,从看似与著名诗人无异的诗歌到足以让人误以为是真人聊天的俚语。尽管风格差异对普通人不明显,但每个人的写作都具有独特的语言特征,如同语言指纹。本研究探讨语言模型是否也有专属指纹。通过风格学与多维语域分析,比较人类与模型在不同语境下的文本。结果发现,模型虽能根据提示调整风格(如维基百科条目或大学论文),但变化范围有限,无法完全模拟人类。具体而言,模型更倾向于使用名词而非动词,表现出与人类不同的语言骨架。人类语言高度依赖时态、体貌、语气等复杂语法维度,而这些可能反映人类独有的思维方式,可作为判断人工智能的试金石。

原文摘要 · Abstract (English)

Large Language Models can emulate different writing styles, ranging from composing poetry that appears indistinguishable from that of famous poets to using slang that can convince people that they are chatting with a human online. While differences in style may not always be visible to the untrained eye, we can generally distinguish the writing of different people, like a linguistic fingerprint. This work examines whether a language model can also be linked to a specific fingerprint. Through stylometric and multidimensional register analyses, we compare human-authored and model-authored texts from different registers. We find that the model can successfully adapt its style depending on whether it is prompted to produce a Wikipedia entry vs. a college essay, but not in a way that makes it indistinguishable from humans. Concretely, the model shows more limited variation when producing outputs in different registers. Our results suggest that the model prefers nouns to verbs, thus showing a distinct linguistic backbone from humans, who tend to anchor language in the highly grammaticalized dimensions of tense, aspect, and mood. It is possible that the more complex domains of grammar reflect a mode of thought unique to humans, thus acting as a litmus test for Artificial Intelligence.

语言指纹AI检测自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。