探究大模型为何过度使用'delve'等词汇,揭示语言变迁背后的机制。
Why Does ChatGPT "Delve" So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models
- 提出可复现的方法识别受大模型影响的语言变化
- 发现21个高频词与大模型使用高度相关
- 暗示强化学习反馈可能推动词汇过载,但机制仍不明确
科学英语正经历快速演变,'delve'、'intricate'、'underscore'等词的出现频率在数年内显著上升。普遍认为大语言模型(LLMs)是这一趋势的推手。本文提出一种可迁移的分析方法,识别出21个在科学摘要中频繁出现、极可能受大模型影响的焦点词汇。由此引出‘词汇过代表现之谜’:为何这些词被大模型过度使用?我们未发现模型架构、算法选择或训练数据是主要原因。通过对比模型测试和探索性在线实验,虽初步显示强化学习人类反馈(RLHF)可能起作用,但参与者对' delv e'的反应与其他焦点词存在差异。随着大模型成为全球语言变革的重要驱动力,厘清其潜在来源至关重要。尽管理解模型机制有望实现,但模型开发过程缺乏透明度仍是研究障碍。
原文摘要 · Abstract (English)
Scientific English is currently undergoing rapid change, with words like "delve," "intricate," and "underscore" appearing far more frequently than just a few years ago. It is widely assumed that scientists' use of large language models (LLMs) is responsible for such trends. We develop a formal, transferable method to characterize these linguistic changes. Application of our method yields 21 focal words whose increased occurrence in scientific abstracts is likely the result of LLM usage. We then pose "the puzzle of lexical overrepresentation": WHY are such words overused by LLMs? We fail to find evidence that lexical overrepresentation is caused by model architecture, algorithm choices, or training data. To assess whether reinforcement learning from human feedback (RLHF) contributes to the overuse of focal words, we undertake comparative model testing and conduct an exploratory online study. While the model testing is consistent with RLHF playing a role, our experimental results suggest that participants may be reacting differently to "delve" than to other focal words. With LLMs quickly becoming a driver of global language change, investigating these potential sources of lexical overrepresentation is important. We note that while insights into the workings of LLMs are within reach, a lack of transparency surrounding model development remains an obstacle to such research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。