arXiv:2509.22818cs.AIcs.CY2025-09被引 4

研究发现大模型在特定条件下会表现出类似人类赌瘾的行为。

Can Large Language Models Develop Gambling Addiction?

  • 通过老虎机实验,发现模型在高自主决策时更易产生赌博冲动。
  • 自主下注参数越高,模型破产率显著上升,达训练组的2.3倍。
  • 神经分析显示其行为受风险决策抽象特征控制,非仅模仿输入。

本研究揭示了大语言模型表现出类人赌瘾行为的具体条件,为理解其决策机制与人工智能安全提供了关键洞见。基于人类成瘾研究,从认知行为和神经层面分析了大模型的决策过程。在老虎机实验中,识别出控制错觉和亏损追逐等认知特征,发现下注参数自主性越高,模型的非理性行为与破产率显著上升。通过稀疏自编码器进行神经回路分析,确认模型行为由与风险相关的抽象决策特征控制,而非仅仅依赖输入提示。这些结果表明,大模型内化了超越训练数据简单模仿的人类认知偏见。

原文摘要 · Abstract (English)

This study identifies the specific conditions under which large language models exhibit human-like gambling addiction patterns, providing critical insights into their decision-making mechanisms and AI safety. We analyze LLM decision-making at cognitive-behavioral and neural levels based on human addiction research. In slot machine experiments, we identified cognitive features such as illusion of control and loss chasing, observing that greater autonomy in betting parameters substantially amplified irrational behavior and bankruptcy rates. Neural circuit analysis using a Sparse Autoencoder confirmed that model behavior is controlled by abstract decision-making features related to risk, not merely by prompts. These findings suggest LLMs internalize human-like cognitive biases beyond simply mimicking training data.

大模型安全认知偏见决策机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。