ChatGPT让人类口语中出现新词汇,表明大模型正影响人类语言演化。
Empirical evidence of Large Language Model's influence on human spoken communication
- 通过分析海量播客语音,发现特定词汇在ChatGPT发布后突然增多
- 实验证明用户短暂对话后会持续使用模型词汇,且难以改变
- 研究揭示大模型已开始反向塑造人类语言,可能引发文化同质化
从印刷术到社交媒体,通信技术的每一次革新都重塑了思想传播方式。由生成式人工智能驱动的聊天机器人构成一种新媒介,其神经表征编码文化模式,并在与数亿人对话中传播。这些模式是否进入人类语言并最终影响文化,是一个根本性问题。尽管完全量化如ChatGPT对人类文化的因果影响极具挑战,但人类口语中的词汇变化可作为早期指标。本研究发现,被ChatGPT频繁生成的词汇(如delve、showcase、boast、intricacies、meticulous)在非脚本口语中显著增加。基于824,634个播客片段、总计737,083小时未剪辑对话的合成控制分析,证实这一变化与ChatGPT发布存在因果关联。一项预注册实验(N = 496)进一步确认,用户简短对话后会持续使用模型词汇,即使经过干扰任务仍能识别,表明其已融入主动词汇库。结果表明,基于人类数据训练的机器正在将自身特征反馈至人类语言,嵌入文化演化的持续过程。这种耦合引发了对语言同质化及少数大型AI提供者潜在大规模文化影响力的新担忧。
原文摘要 · Abstract (English)
From the printing press to social media, innovations in communication technology have repeatedly reshaped how ideas spread through human culture. Chatbots powered by generative artificial intelligence constitute a new medium, encoding cultural patterns in their neural representations and disseminating them in conversations with hundreds of millions of people. Whether these patterns transmit into human language, and ultimately shape human culture, is a fundamental question. While fully quantifying the causal impact of a chatbot like ChatGPT on human culture is challenging, lexical shifts in human spoken communication may offer an early indicator. Here we show that words preferentially generated by ChatGPT, such as delve, showcase, boast, intricacies and meticulous, increased abruptly in spontaneous human speech. A synthetic-control analysis of 737,083 hours of conversation from 824,634 podcast episodes, screened for unscripted speech, causally links this shift to ChatGPT's release. The measurable influence on spontaneous speech suggests that humans internalize the lexical choices of large language models (LLMs). A preregistered experiment (N = 496) confirms they do, as a brief chatbot interaction led participants to adopt its words as their own, persisting past a distractor task and confirmed in forced lexical choice, indicating entrenchment in the active vocabulary. Together these results show that machines trained on human data now feed their own traits back into human language, integrating LLMs into the ongoing processes of cultural evolution.. This coupling raises concerns about linguistic homogenization and the capacity of a few major AI providers for latent cultural influence at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。