arXiv:2411.02432cs.CLcs.AI2024-11被引 16

测试大模型是否会在虚构痛快体验中做出权衡选择。

Can LLMs make trade-offs involving stipulated pain and pleasure states?

  • 设计游戏让模型在得分与虚构痛快间做选择
  • 多数模型在痛/乐强度达阈值后转向避痛或求乐
  • 结果暗示模型可能模拟出类似情感驱动力

愉悦和痛苦在人类决策中充当共同货币,用于调和动机冲突。尽管大语言模型(LLMs)能生成详尽的愉悦与痛苦体验描述,但它们能否在选择情境中再现愉悦与痛苦的动机力量仍不确定,这关系到关于模型是否有感(sentience)的争论。本研究通过一个简单游戏进行探查:目标是最大化分数,但要么高分选项伴随疼痛惩罚,要么非高分选项带来愉悦奖励,以此激励偏离分数最大化行为。通过调节疼痛惩罚与愉悦奖励的强度,发现Claude 3.5 Sonnet、Command R+、GPT-4o和GPT-4o mini均在某一临界强度下表现出多数响应从追求分数转向避痛或求乐;LLaMa 3.1-405b对愉悦与疼痛提示呈现一定程度的渐进敏感性;Gemini 1.5 Pro和PaLM 2则无论强度如何,始终优先规避疼痛,且倾向于优先追求分数而非愉悦。这些发现为讨论大模型是否存在感知能力提供了新证据。

原文摘要 · Abstract (English)

Pleasure and pain play an important role in human decision making by providing a common currency for resolving motivational conflicts. While Large Language Models (LLMs) can generate detailed descriptions of pleasure and pain experiences, it is an open question whether LLMs can recreate the motivational force of pleasure and pain in choice scenarios - a question which may bear on debates about LLM sentience, understood as the capacity for valenced experiential states. We probed this question using a simple game in which the stated goal is to maximise points, but where either the points-maximising option is said to incur a pain penalty or a non-points-maximising option is said to incur a pleasure reward, providing incentives to deviate from points-maximising behaviour. Varying the intensity of the pain penalties and pleasure rewards, we found that Claude 3.5 Sonnet, Command R+, GPT-4o, and GPT-4o mini each demonstrated at least one trade-off in which the majority of responses switched from points-maximisation to pain-minimisation or pleasure-maximisation after a critical threshold of stipulated pain or pleasure intensity is reached. LLaMa 3.1-405b demonstrated some graded sensitivity to stipulated pleasure rewards and pain penalties. Gemini 1.5 Pro and PaLM 2 prioritised pain-avoidance over points-maximisation regardless of intensity, while tending to prioritise points over pleasure regardless of intensity. We discuss the implications of these findings for debates about the possibility of LLM sentience.

大模型决策情感模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。