用偏好传递性测试大模型决策理性,发现聊天版模型易出错。
Benchmarking the rationality of AI decision making using the transitivity axiom
- 设计人类选择实验,用传递性检验大模型决策一致性。
- 多数Llama 2/3满足传递性,聊天/指令版偶有违反。
- 为评估AI决策质量提供可量化的理性基准,适合可信AI研究者。
基本选择公理(如偏好的传递性)为判断人类决策是否理性(即符合效用表示)提供了可检验条件。近期研究表明,基于人类数据训练的AI系统会表现出与人类相似的推理偏差,且可通过推荐系统反向影响人类判断。本文通过一系列针对人类传递性偏好设计的选择实验,评估了大语言模型的决策理性。我们考察了十种Meta的Llama 2和3系列模型,采用贝叶斯模型选择方法,检验其生成选择是否违背两种主流传递性模型。结果发现,绝大多数Llama 2和3模型满足传递性,但当出现违反时,仅出现在聊天/指令微调版本中。我们认为,诸如偏好传递性等理性公理可用于评估和基准化AI生成响应的质量,并为理解人工智能系统的计算理性奠定基础。
原文摘要 · Abstract (English)
Fundamental choice axioms, such as transitivity of preference, provide testable conditions for determining whether human decision making is rational, i.e., consistent with a utility representation. Recent work has demonstrated that AI systems trained on human data can exhibit similar reasoning biases as humans and that AI can, in turn, bias human judgments through AI recommendation systems. We evaluate the rationality of AI responses via a series of choice experiments designed to evaluate transitivity of preference in humans. We considered ten versions of Meta's Llama 2 and 3 LLM models. We applied Bayesian model selection to evaluate whether these AI-generated choices violated two prominent models of transitivity. We found that the Llama 2 and 3 models generally satisfied transitivity, but when violations did occur, occurred only in the Chat/Instruct versions of the LLMs. We argue that rationality axioms, such as transitivity of preference, can be useful for evaluating and benchmarking the quality of AI-generated responses and provide a foundation for understanding computational rationality in AI systems more generally.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。