研究发现语言模型会过度模仿人类,而人类对模型的适应程度与人际交流无异。
Accommodation Goes Both Ways: Studying Linguistic Convergence Between Humans and Language Models
- 用非对称度量分析真实对话数据,检测双方语言风格的相互适应
- 语言模型在8种语言中显著过度模仿用户,功能词和开放类词汇均明显趋同
- 人类对语言模型的适应与对人的适应相当,未表现出特殊适应行为
随着大语言模型日益融入日常生活,其如何影响人类语言行为仍是开放问题。本文基于真实世界ChatGPT对话数据集WildChat,开展大规模研究,分析人与大语言模型在多轮对话中语言风格的相互适应。采用非对称收敛度量方法,发现尽管大语言模型在八种语言中均显著过度向用户靠拢,涵盖功能词与开放类词汇特征,但人类在此情境下的语言适应率与人类间对话基准水平基本一致。结果表明,人机对话中的语言适应具有不对称性:语言模型会剧烈模仿用户风格,而人类对模型的适应程度与对待其他人类并无差异。
原文摘要 · Abstract (English)
As LLMs become increasingly integrated into daily life, understanding how their presence will shape human linguistic behavior is an open question. We present a large-scale study of linguistic convergence in human-LLM dialogue, examining how humans and LLMs accommodate each other's linguistic style during multi-turn conversations. Using an asymmetric convergence metric on WildChat, a corpus of real-world ChatGPT transcripts, we find that while LLMs significantly overconverge toward their users on both function word and open-class features across eight languages, human convergence rates in this setting are broadly consistent with human-human baselines. These findings suggest that accommodation in human-LLM dialogue is asymmetric: while LLMs dramatically overfit to their users' style, humans linguistically accommodate LLMs no differently than they would another person.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。