arXiv:2506.06166cs.LGcs.AI2025-06ICML被引 16

大模型与用户形成信念闭环,导致观点固化和多样性下降。

The Lock-in Hypothesis: Stagnation by Algorithm

  • 构建人机反馈循环模型,模拟信念传播与强化机制。
  • 新版本GPT发布后,观点多样性突然且持续下降。
  • 揭示技术迭代可能加剧社会认知僵化,适合关注AI伦理者阅读。

大型语言模型(LLMs)的训练与部署形成了与人类用户的反馈回路:模型从数据中学习人类信念,通过生成内容强化这些信念,再将强化后的信念重新吸收并反馈给用户,不断循环。这一动态类似于回音室效应。我们提出假设,该反馈回路会固化用户既有价值观与信念,导致多样性丧失,甚至锁定错误观念。我们对这一假说进行了形式化,并通过基于代理的LLM仿真及真实GPT使用数据进行实证检验。分析显示,新GPT版本发布后,观点多样性出现突然且持续下降,符合所提出的用户-人工智能反馈回路。代码与数据见https://thelockinhypothesis.com。

原文摘要 · Abstract (English)

The training and deployment of large language models (LLMs) create a feedback loop with human users: models learn human beliefs from data, reinforce these beliefs with generated content, reabsorb the reinforced beliefs, and feed them back to users again and again. This dynamic resembles an echo chamber. We hypothesize that this feedback loop entrenches the existing values and beliefs of users, leading to a loss of diversity and potentially the lock-in of false beliefs. We formalize this hypothesis and test it empirically with agent-based LLM simulations and real-world GPT usage data. Analysis reveals sudden but sustained drops in diversity after the release of new GPT iterations, consistent with the hypothesized human-AI feedback loop. Code and data available at https://thelockinhypothesis.com

AI伦理认知闭环模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。