arXiv:2510.05136cs.CLcs.AI2025-10综述被引 10

系统梳理AI生成文本的语言特征,揭示其形式化与重复性问题。

Linguistic Characteristics of AI-Generated Text: A Survey

  • 按语言层级、模型、语体等维度整合现有研究
  • 发现AI文本更正式、词汇多样性低且易重复
  • 适合关注生成文本质量与评估的研究者

大型语言模型(LLMs)在教育、医疗和科研等领域已成为文本自动生成的有力工具。随着其应用日益广泛,研究AI生成文本的语言特征变得至关重要,这关系到语料库语言学、计算语言学及自然语言处理等多个领域。尽管已有诸多观察,但尚缺乏对现有成果的全面综述。本文旨在系统梳理相关研究,从语言描述层次、模型类型、文本体裁、语言种类及提示方式等维度进行分类。结果显示,AI生成文本普遍呈现更正式、非个人化的风格,表现为名词、限定词和介词使用增多,形容词与副词减少;同时词汇多样性降低、词汇量小且存在重复现象。然而,当前研究仍集中于英语和GPT系列模型,跨语言与跨模型研究不足,且多数未考虑提示词敏感性,未来需采用多提示策略开展更深入研究。

原文摘要 · Abstract (English)

Large language models (LLMs) are solidifying their position in the modern world as effective tools for the automatic generation of text. Their use is quickly becoming commonplace in fields such as education, healthcare, and scientific research. There is a growing need to study the linguistic features present in AI-generated text, as the increasing presence of such texts has profound implications in various disciplines such as corpus linguistics, computational linguistics, and natural language processing. Many observations have already been made, however a broader synthesis of the findings made so far is required to provide a better understanding of the topic. The present survey paper aims to provide such a synthesis of extant research. We categorize the existing works along several dimensions, including the levels of linguistic description, the models included, the genres analyzed, the languages analyzed, and the approach to prompting. Additionally, the same scheme is used to present the findings made so far and expose the current trends followed by researchers. Among the most-often reported findings is the observation that AI-generated text is more likely to contain a more formal and impersonal style, signaled by the increased presence of nouns, determiners, and adpositions and the lower reliance on adjectives and adverbs. AI-generated text is also more likely to feature a lower lexical diversity, a smaller vocabulary size, and repetitive text. Current research, however, remains heavily concentrated on English data and mostly on text generated by the GPT model family, highlighting the need for broader cross-linguistic and cross-model investigation. In most cases authors also fail to address the issue of prompt sensitivity, leaving much room for future studies that employ multiple prompt wordings in the text generation phase.

语言特征AI生成文本大模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。