arXiv:2506.18600cs.CLcs.GT2025-06被引 2

反驳数据泄露质疑,证明大模型群体可研究真实涌现行为

Reply to "Emergent LLM behaviors are observationally equivalent to data leakage"

  • 通过反例说明模型自组织与涌现现象仍可研究
  • 实证发现社会惯例在模型群体中自然形成
  • 适合关注多智能体系统与模型涌现的研究者

模拟大型语言模型(LLM)群体时,数据污染可能引发训练数据意外影响结果的担忧。尽管该问题重要且可能阻碍某些多智能体实验,但并不妨碍研究真实涌现的动力学。针对Barrie和Törnberg对Flint Ashery等人研究结果的批评,本文澄清:自组织及依赖模型的涌现行为可在LLM群体中研究,并以社会惯例的实证观察为例,展示了此类动态的可验证性。

原文摘要 · Abstract (English)

A potential concern when simulating populations of large language models (LLMs) is data contamination, i.e. the possibility that training data may shape outcomes in unintended ways. While this concern is important and may hinder certain experiments with multi-agent models, it does not preclude the study of genuinely emergent dynamics in LLM populations. The recent critique by Barrie and Törnberg [1] of the results of Flint Ashery et al. [2] offers an opportunity to clarify that self-organisation and model-dependent emergent dynamics can be studied in LLM populations, highlighting how such dynamics have been empirically observed in the specific case of social conventions.

大模型群体涌现行为多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。