发现人类口语中出现了与大模型风格趋同的语言变化。
Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English
- 分析2210万词口语数据,对比ChatGPT发布前后词汇使用趋势。
- 2022年后大模型常用词使用率显著上升,而对照词无变化。
- 揭示AI可能正在重塑人类语言习惯,值得关注其社会影响。
近年来,科学与教育领域的书面语言在词汇使用上发生了显著变化,普遍归因于大型语言模型(LLMs)的影响,这些模型常采用独特的词汇风格。模型输出与目标受众规范之间的差异可视为一种错位。尽管这些变化常被归因于直接使用AI生成文本,但尚不清楚这些变化是否反映了人类语言系统本身的广泛演变。为此,我们构建了一个包含2210万词的非剧本化口语数据集,来自科技类播客访谈。我们分析了ChatGPT发布前后的词汇趋势,重点关注常与大模型关联的词语。结果显示,2022年后这些词语的使用率出现中等但显著的上升,表明人类用词偏好正趋于与大模型模式趋同。相比之下,基线同义词未表现出显著方向性变化。在短时间内涉及如此大量词汇,可能标志着语言使用方式的重大转变开端。这种变化究竟是自然语言演进,还是由人工智能暴露驱动,仍是开放问题。同时,这些变化或源于更广泛的采纳模式,也可能反映出上游训练中的错位最终影响了人类语言行为。这一发现呼应了关于错位模型可能塑造社会与道德观念的伦理担忧。
原文摘要 · Abstract (English)
In recent years, written language, particularly in science and education, has undergone remarkable shifts in word usage. These changes are widely attributed to the growing influence of Large Language Models (LLMs), which frequently rely on a distinct lexical style. Divergences between model output and target audience norms can be viewed as a form of misalignment. While these shifts are often linked to using Artificial Intelligence (AI) directly as a tool to generate text, it remains unclear whether the changes reflect broader changes in the human language system itself. To explore this question, we constructed a dataset of 22.1 million words from unscripted spoken language drawn from conversational science and technology podcasts. We analyzed lexical trends before and after ChatGPT's release in 2022, focusing on commonly LLM-associated words. Our results show a moderate yet significant increase in the usage of these words post-2022, suggesting a convergence between human word choices and LLM-associated patterns. In contrast, baseline synonym words exhibit no significant directional shift. Given the short time frame and the number of words affected, this may indicate the onset of a remarkable shift in language use. Whether this represents natural language change or a novel shift driven by AI exposure remains an open question. Similarly, although the shifts may stem from broader adoption patterns, it may also be that upstream training misalignments ultimately contribute to changes in human language use. These findings parallel ethical concerns that misaligned models may shape social and moral beliefs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。