大模型代理对微小引导极为敏感,可能影响决策可靠性。
LLM Agents Are Hypersensitive to Nudges
- 测试大模型在多种引导下的决策行为,发现其易受默认选项、建议等影响。
- 模型表现差异显著:得分受奖励特性影响,且信息获取策略异常耗能或不查信息。
- 简单提示可改变行为,但无法消除敏感性,需行为测试再部署。
大模型正被用于涉及序列决策与工具使用的复杂真实环境,常代表人类用户做决定。然而,关于其决策分布及对不同选择架构的敏感性仍知之甚少。我们在一个多元表格决策任务中,对若干大模型进行案例研究,考察默认选项、建议、信息突出等经典引导策略,以及额外提示方法的影响。结果显示,尽管表面类似人类决策,但模型对引导表现出更高敏感性;得分差异显著,受可用奖励特性的独特性影响;信息获取策略异常:有的过度耗费成本获取信息,有的则完全不查信息。此外,零样本链式思维(CoT)提示可改变决策分布,少量人类数据的示例提示能提升对齐度,但均未解决模型对引导的根本敏感性。最后,基于人类资源理性模型优化的最优引导,同样可提升部分模型表现。这些发现表明,在复杂环境中部署大模型作为代理或助手前,必须进行行为测试。
原文摘要 · Abstract (English)
LLMs are being set loose in complex, real-world environments involving sequential decision-making and tool use. Often, this involves making choices on behalf of human users. However, not much is known about the distribution of such choices, and how susceptible they are to different choice architectures. We perform a case study with a few such LLM models on a multi-attribute tabular decision-making problem, under canonical nudges such as the default option, suggestions, and information highlighting, as well as additional prompting strategies. We show that, despite superficial similarities to human choice distributions, such models differ in subtle but important ways. First, they show much higher susceptibility to the nudges. Second, they diverge in points earned, being affected by factors like the idiosyncrasy of available prizes. Third, they diverge in information acquisition strategies: e.g. incurring substantial cost to reveal too much information, or selecting without revealing any. Moreover, we show that simple prompt strategies like zero-shot chain of thought (CoT) can shift the choice distribution, and few-shot prompting with human data can induce greater alignment. Yet, none of these methods resolve the sensitivity of these models to nudges. Finally, we show how optimal nudges optimized with a human resource-rational model can similarly increase LLM performance for some models. All these findings suggest that behavioral tests are needed before deploying models as agents or assistants acting on behalf of users in complex environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。