大模型输出有独特语言风格,像人类的个人语言习惯。
Beyond "AI Language": The case for the idiolectal nature of LLM output
- 用风格分析方法识别不同大模型的语言特征
- 2026年模型中缩略语使用量每百万词超3万,差异显著
- 适合研究语言演变、文本检测和语言学分析的人看
尽管大语言模型输出常被视作一种统称的'人工智能语言',本文主张其同时具有独特的、模型特异的语言签名,类似人类的个人语言习惯。我们分析了两个关于社会议题的LLM生成文本数据集:2024年六款模型的语料库(Improta et al. 2024)以及用相同提示生成的2026年六款现代模型的新语料库。通过计算描述符与风格学主成分分析,发现2024与2026两代模型间存在风格变迁,且每个模型保持独特语言特征。这一多层互动在缩略语频率上体现明显——同一届模型中,每百万词使用量从超过1200到超过3万不等。结论指出,将LLM输出视为具有个体性特征,有助于推动语言变异与变化研究、生成文本检测、司法语言学及基于使用的语言研究。
原文摘要 · Abstract (English)
While large language model outputs are frequently analysed as a collective super variety termed "AI language," this chapter argues that this perspective coexists with distinct, model-specific linguistic signatures akin to human idiolects. We analyse two datasets of LLM-generated texts on societal topics: a 2024 corpus of six models (Improta et al. 2024) and a newly generated 2026 corpus using the same prompts featuring six contemporary models. Our findings, utilising computational descriptors and stylometric principal component analysis reveal a generational shift between the style of the 2024 and 2026 cohorts, while demonstrating that each individual model maintains a unique linguistic profile. This multi-layered interplay is illustrated by contraction frequencies, which vary from over 1,200 to over 30,000 per million words within the same cohort of models (2026). Ultimately, we conclude that treating LLM output as idiolectal in nature provides a valuable framework with potential implications for research on variation and change, LLM-generated text detection, forensic linguistics and usage-based approaches to language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。