arXiv:2511.08721cs.LGcs.AI2025-11被引 2

研究大模型在分钱游戏中的行为,发现提示词能显著影响其公平性表现。

Benevolent Dictators? On LLM Agent Behavior in Dictator Games

  • 设计新框架分析不同系统提示对大模型行为的影响
  • 大模型多表现出强烈公平倾向,但结果受提示微调影响大
  • 通过语言特征分析揭示模型决策背后的逻辑

在行为科学中,独裁者博弈被用于评估参与者对公平或自利的偏好。在该博弈中,一名独裁者单方面决定如何分配固定金额给自身与另一方。尽管已有研究探索基于大语言模型(LLMs)的智能体在不同人格设定下的行为模式,但这些研究往往忽视了系统提示——即塑造模型行为的基础指令——的作用,且未考虑结果对提示微小变化的敏感性。一个稳健的基线对于研究大模型复杂的类人行为至关重要。为此,我们提出大模型智能体行为研究框架(LLM-ABS),旨在:(i) 探索不同系统提示如何影响模型行为;(ii) 通过中性提示变体获得更可靠的智能体偏好洞察;(iii) 分析大模型在开放指令下的语言特征,以理解其行为背后的原因。我们发现,模型通常表现出强烈的公平偏好,且系统提示对其行为有显著影响。从语言角度,模型回应方式存在差异。尽管提示敏感性仍是挑战,但本框架为大模型行为研究提供了稳健基础。代码已公开于 https://github.com/andreaseinwiller/LLM-ABS。

原文摘要 · Abstract (English)

In behavioral sciences, experiments such as the ultimatum game are conducted to assess preferences for fairness or self-interest of study participants. In the dictator game, a simplified version of the ultimatum game where only one of two players makes a single decision, the dictator unilaterally decides how to split a fixed sum of money between themselves and the other player. Although recent studies have explored behavioral patterns of AI agents based on Large Language Models (LLMs) instructed to adopt different personas, we question the robustness of these results. In particular, many of these studies overlook the role of the system prompt - the underlying instructions that shape the model's behavior - and do not account for how sensitive results can be to slight changes in prompts. However, a robust baseline is essential when studying highly complex behavioral aspects of LLMs. To overcome previous limitations, we propose the LLM agent behavior study (LLM-ABS) framework to (i) explore how different system prompts influence model behavior, (ii) get more reliable insights into agent preferences by using neutral prompt variations, and (iii) analyze linguistic features in responses to open-ended instructions by LLM agents to better understand the reasoning behind their behavior. We found that agents often exhibit a strong preference for fairness, as well as a significant impact of the system prompt on their behavior. From a linguistic perspective, we identify that models express their responses differently. Although prompt sensitivity remains a persistent challenge, our proposed framework demonstrates a robust foundation for LLM agent behavior studies. Our code artifacts are available at https://github.com/andreaseinwiller/LLM-ABS.

大模型行为提示工程公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。