arXiv:2505.15011cs.AI2025-05被引 1

通过声誉加权奖励,让智能体同时遵守明文规则与社会规范。

HAVA: Hybrid Approach to Value-Alignment through Reward Weighing for Reinforcement Learning

  • 用声誉量化智能体对规则与规范的遵守程度,动态调整奖励。
  • 在交通场景中验证,同时考虑显性与隐性规范可提升价值对齐效果。
  • 适合需要兼顾法律合规与社会行为的强化学习应用。

我们的社会由一系列规范构成,共同体现我们珍视的价值,如安全、公平和可信。价值对齐的目标是使智能体不仅能完成任务,还能通过行为促进这些价值。许多规范以法律或规则形式明确表达(法律/安全规范),而更多则为非正式的社会规范。此外,这些规范的表示方式也不同:安全规范通常以逻辑语言显式表达,而社会规范则通常隐藏在神经网络参数空间中。现有研究缺乏将这些不同表示的规范整合到单一算法中的方法。本文提出一种新方法,将各类规范融入强化学习过程。该方法监控智能体对规范的遵守情况,并将其总结为“声誉”这一量度,用以加权接收到的奖励,从而激励智能体实现价值对齐。我们在连续状态空间的交通问题中进行了一系列实验,验证了显性与隐性规范的重要性,并展示了本方法如何找到价值对齐策略。此外,消融实验表明,结合两类规范优于单独使用其中任何一类。

原文摘要 · Abstract (English)

Our society is governed by a set of norms which together bring about the values we cherish such as safety, fairness or trustworthiness. The goal of value-alignment is to create agents that not only do their tasks but through their behaviours also promote these values. Many of the norms are written as laws or rules (legal / safety norms) but even more remain unwritten (social norms). Furthermore, the techniques used to represent these norms also differ. Safety / legal norms are often represented explicitly, for example, in some logical language while social norms are typically learned and remain hidden in the parameter space of a neural network. There is a lack of approaches in the literature that could combine these various norm representations into a single algorithm. We propose a novel method that integrates these norms into the reinforcement learning process. Our method monitors the agent's compliance with the given norms and summarizes it in a quantity we call the agent's reputation. This quantity is used to weigh the received rewards to motivate the agent to become value-aligned. We carry out a series of experiments including a continuous state space traffic problem to demonstrate the importance of the written and unwritten norms and show how our method can find the value-aligned policies. Furthermore, we carry out ablations to demonstrate why it is better to combine these two groups of norms rather than using either separately.

强化学习价值对齐声誉机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。