通过语言特征分析,揭示人类与大模型文本的差异与趋同。
Linguistic and Embedding-Based Profiling of Texts generated by Humans and Large Language Models
- 从形态、句法、语义多层面提取语言特征进行对比分析。
- 人类文本句法更简单,语义更丰富;模型文本随时间趋于同质化。
- 新模型生成文本风格趋同,适合内容安全与检测研究者参考。
大语言模型(LLMs)生成文本的能力迅速提升,使其与人类写作的界限日益模糊。本文聚焦于使用多种语言特征(如依存长度、情感性)对来自11个不同模型、覆盖8个领域的生成文本与人类文本进行刻画。分析发现,人类文本普遍具有更简单的句法结构和更丰富的语义内容。我们还计算了特征在模型与领域间的变异性:尽管两类文本在各领域均呈现风格多样性,但人类文本在特征上表现出更大波动。进一步通过风格嵌入验证,结果显示新发布模型生成的文本风格趋于一致,表明机器生成文本正出现同质化趋势。
原文摘要 · Abstract (English)
The rapid advancements in large language models (LLMs) have significantly improved their ability to generate natural language, making texts generated by LLMs increasingly indistinguishable from human-written texts. While recent research has primarily focused on using LLMs to classify text as either human-written or machine-generated texts, our study focuses on characterizing these texts using a set of linguistic features across different linguistic levels such as morphology, syntax, and semantics. We select a dataset of human-written and machine-generated texts spanning 8 domains and produced by 11 different LLMs. We calculate different linguistic features such as dependency length and emotionality, and we use them for characterizing human-written and machine-generated texts along with different sampling strategies, repetition controls, and model release dates. Our statistical analysis reveals that human-written texts tend to exhibit simpler syntactic structures and more diverse semantic content. Furthermore, we calculate the variability of our set of features across models and domains. Both human- and machine-generated texts show stylistic diversity across domains, with human-written texts displaying greater variation in our features. Finally, we apply style embeddings to further test variability among human-written and machine-generated texts. Notably, newer models output text that is similarly variable, pointing to a homogenization of machine-generated texts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。