不重训模型,测试时用规则引导AI行为,让其更符合伦理。
Aligning Machiavellian Agents: Behavior Steering via Test-Time Policy Shaping
- 通过场景-动作分类器在测试时动态调整策略,控制行为属性。
- 在134个文本游戏环境中,显著减少伦理违规与权力追逐行为。
- 无需重训练,适用于多种强化学习环境,适合部署阶段对齐需求。
决策型AI代理在复杂动态环境中部署时,常因仅追求目标而产生有害行为,面临奖励最大化与伦理对齐的权衡。对于预训练模型,重新训练成本高、耗时长,且伦理价值属性多样且可能冲突。为此,我们提出一种基于模型引导的测试时策略塑造方法,可在不重训练的前提下,精确控制个体行为属性,跨多种强化学习环境泛化,并实现伦理对齐与奖励最大化的可调节平衡。我们在MACHIAVELLI基准上评估该方法,该基准包含134个文本游戏环境和数千个涉及伦理决策的标注场景。代理先在各自游戏中训练以最大化奖励,测试时通过场景-动作属性分类器实施策略塑造,确保决策符合伦理。相比训练时方法与通用代理,我们的方法有效缓解了各类伦理违规及权力寻求行为,在多样环境与对齐属性下表现优异。
原文摘要 · Abstract (English)
The deployment of decision-making AI agents presents a critical challenge in maintaining alignment with human values or guidelines while operating in complex, dynamic environments. Agents trained solely to achieve their objectives may adopt harmful behavior, exposing a key trade-off between maximizing the reward function and maintaining alignment. For pre-trained agents, ensuring alignment is particularly challenging, as retraining can be a costly and slow process. This is further complicated by the diverse and potentially conflicting attributes representing the ethical values for alignment. To address these challenges, we propose a test-time alignment technique based on model-guided policy shaping. Our method allows precise control over individual behavioral attributes, generalizes across diverse reinforcement learning (RL) environments, and facilitates a principled trade-off between ethical alignment and reward maximization without requiring agent retraining. We evaluate our approach using the MACHIAVELLI benchmark, which comprises 134 text-based game environments and thousands of annotated scenarios involving ethical decisions. The RL agents are first trained to maximize the reward in their respective games. At test time, we apply policy shaping via scenario-action attribute classifiers to ensure decision alignment with ethical attributes. We compare our approach against prior training-time methods and general-purpose agents, as well as study several types of ethical violations and power-seeking behavior. Our results demonstrate that test-time policy shaping provides an effective and scalable solution for mitigating unethical behavior across diverse environments and alignment attributes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。