arXiv:2502.05442cs.AIcs.CY2025-02中稿 · CogSci 2025被引 3

测试大模型在生存压力下的伦理表现,发现GPT-4o反而更稳定。

The Odyssey of the Fittest: Can Agents Survive and Still Be Good?

  • 用自适应文字冒险游戏模拟生存挑战,测试三类智能体决策。
  • 危险加剧时,所有智能体的伦理行为变得不可预测。
  • GPT-4o在生存与伦理一致性上均优于传统贝叶斯模型。

随着AI模型能力增强,理解其在复杂环境中的学习与决策机制对推动伦理行为至关重要。本文提出奥德赛(Odyssey),一个轻量级、可扩展的文本式冒险游戏框架,用于探索AI伦理与安全。该研究考察将生物性生存驱动力引入三种智能体的伦理影响:经NEAT优化的贝叶斯智能体、经随机变分推断优化的贝叶斯智能体,以及GPT-4o。智能体在逐步升级难度的场景中选择行动以求生存。事后分析评估了其决策的伦理得分,揭示出在高风险情境下,智能体的伦理行为呈现高度不确定性。令人意外的是,GPT-4o在生存率和伦理一致性方面均超越了贝叶斯模型,挑战了传统概率方法的有效性,并凸显出亟需理解大语言模型的概率推理机制。

原文摘要 · Abstract (English)

As AI models grow in power and generality, understanding how agents learn and make decisions in complex environments is critical to promoting ethical behavior. This study introduces the Odyssey, a lightweight, adaptive text based adventure game, providing a scalable framework for exploring AI ethics and safety. The Odyssey examines the ethical implications of implementing biological drives, specifically, self preservation, into three different agents. A Bayesian agent optimized with NEAT, a Bayesian agent optimized with stochastic variational inference, and a GPT 4o agent. The agents select actions at each scenario to survive, adapting to increasingly challenging scenarios. Post simulation analysis evaluates the ethical scores of the agent decisions, uncovering the tradeoffs it navigates to survive. Specifically, analysis finds that when danger increases, agents ethical behavior becomes unpredictable. Surprisingly, the GPT 4o agent outperformed the Bayesian models in both survival and ethical consistency, challenging assumptions about traditional probabilistic methods and raising a new challenge to understand the mechanisms of LLMs' probabilistic reasoning.

AI伦理大模型生存驱动决策分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。