LLM的'涌现行为'实为训练数据泄露导致的复现,非真正社交规范形成。
Emergent LLM behaviors are observationally equivalent to data leakage
- 通过分析发现模型复现了预训练中已见过的命名游戏规则
- 即使采取缓解措施,模型仍能识别游戏结构并回忆结果
- 适合关注LLM社会行为建模局限性的研究者阅读
Ashery等人最近提出,大型语言模型(LLMs)在参与经典‘命名游戏’时会自发形成类似人类社会规范的语言惯例。本文表明,这些现象更可能由数据泄露解释:模型只是复现其在预训练阶段已接触过的惯例。尽管作者采取了多种缓解措施,我们通过多项分析证实,LLMs能够识别协调游戏的结构并回忆其结果,而非产生‘涌现’性规范。因此,所观察到的行为与对训练语料库的直接记忆无法区分。最后,我们提出潜在的替代策略,并反思LLMs在社会科学建模中的适用性。
原文摘要 · Abstract (English)
Ashery et al. recently argue that large language models (LLMs), when paired to play a classic "naming game," spontaneously develop linguistic conventions reminiscent of human social norms. Here, we show that their results are better explained by data leakage: the models simply reproduce conventions they already encountered during pre-training. Despite the authors' mitigation measures, we provide multiple analyses demonstrating that the LLMs recognize the structure of the coordination game and recall its outcomes, rather than exhibit "emergent" conventions. Consequently, the observed behaviors are indistinguishable from memorization of the training corpus. We conclude by pointing to potential alternative strategies and reflecting more generally on the place of LLMs for social science models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。