用经济理论指导大模型行为,让智能体更理性或更道德。
Aligning Large Language Model Agents with Rational and Moral Preferences: A Supervised Fine-Tuning Approach
- 基于经济学理论生成最优策略,监督微调模型行为
- 微调后模型在博弈中表现出明显不同的定价与合作模式
- 适合研究AI对齐、伦理决策和市场模拟的学者
随着大语言模型(LLMs)越来越多地作为市场与组织中的自主代理,其在策略环境中的行为具有显著的经济影响。我们发现,未经调整的LLM代理在经典经济博弈中存在系统性偏差,表现为过度合作且对激励反应不足。为此,我们提出一种监督微调方法,使其行为符合明确的经济偏好。具体而言,我们基于两种理想化效用模型——追求自身利益最大化的homo economicus,以及包含康德普遍性的homo moralis——生成最优策略,并利用这些策略所隐含的推理过程引导微调。在小规模、理论驱动的合成数据上进行微调,可引发持久且可解释的战略行为转变。在道德困境与重复双寡头定价等场景中,不同偏好对齐的代理产生系统性差异的均衡结果与价格动态。这些结果将多智能体系统的AI对齐问题转化为目标设计问题,并展示了经济理论如何指导具有战略一致性的智能体设计。
原文摘要 · Abstract (English)
As large language models (LLMs) increasingly act as autonomous agents in markets and organizations, their behavior in strategic environments becomes economically consequential. We document that off-the-shelf LLM agents exhibit systematic deviations from payoff-sensitive behavior in canonical economic games, including excessive cooperation and limited responsiveness to incentives. We introduce a supervised fine-tuning approach that aligns agent behavior with explicit economic preferences. Specifically, we generate optimal strategies under two stylized utility specifications, homo economicus, which maximizes self-interest, and homo moralis, which incorporates Kantian universalizability, and use these utility-implied reasoning and strategies to guide fine-tuning. Fine-tuning on a small, theory-driven synthetic dataset induces persistent and interpretable shifts in strategic behavior. In applications to moral dilemmas and repeated duopoly pricing, agents aligned to different preference structures produce systematically distinct equilibrium outcomes and pricing dynamics. These results frame AI alignment in multi-agent settings as an objective-design problem and illustrate how economic theory can guide the design of strategically coherent AI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。