用显式道德奖励训练大模型代理,让其行为更符合人类价值观。
Moral Alignment for LLM Agents
- 设计基于道德哲学的内在奖励函数,直接编码人类核心价值。
- 在重复囚徒困境中,道德对齐使代理学会合作并放弃自私策略。
- 该方法可泛化到多种博弈环境,比传统反馈对齐更透明高效。
基于预训练大语言模型(LLM)的决策代理正被广泛部署于人类活动各领域。随着其能力向通用化发展,透明度下降而影响力上升,因此需有效对齐人类价值观。现有对齐方法多依赖人类偏好数据(如RLHF或DPO),但价值隐含且不透明。本文提出一种新方法:使用内在奖励显式、透明地编码核心人类价值观,用于强化学习微调基础代理模型。我们基于义务论与功利主义哲学框架,在重复囚徒困境(IPD)环境中量化道德奖励,评估代理的行为与后果。实验表明,通过道德微调,代理能主动摒弃先前习得的自私策略;且某些道德策略在多种矩阵博弈中具有泛化能力。结果证明,基于内在奖励的微调是一种有前景的通用对齐方案,可能成为当前主流对齐技术的更透明、低成本替代。
原文摘要 · Abstract (English)
Decision-making agents based on pre-trained Large Language Models (LLMs) are increasingly being deployed across various domains of human activity. While their applications are currently rather specialized, several research efforts are underway to develop more generalist agents. As LLM-based systems become more agentic, their influence on human activity will grow and their transparency will decrease. Consequently, developing effective methods for aligning them to human values is vital. The prevailing practice in alignment often relies on human preference data (e.g., in RLHF or DPO), in which values are implicit, opaque and are essentially deduced from relative preferences over different model outputs. In this work, instead of relying on human feedback, we introduce the design of reward functions that explicitly and transparently encode core human values for Reinforcement Learning-based fine-tuning of foundation agent models. Specifically, we use intrinsic rewards for the moral alignment of LLM agents. We evaluate our approach using the traditional philosophical frameworks of Deontological Ethics and Utilitarianism, quantifying moral rewards for agents in terms of actions and consequences on the Iterated Prisoner's Dilemma (IPD) environment. We also show how moral fine-tuning can be deployed to enable an agent to unlearn a previously developed selfish strategy. Finally, we find that certain moral strategies learned on the IPD game generalize to several other matrix game environments. In summary, we demonstrate that fine-tuning with intrinsic rewards is a promising general solution for aligning LLM agents to human values, and it might represent a more transparent and cost-effective alternative to currently predominant alignment techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。