用人类偏好训练自动驾驶伦理决策,发现人言行不一。
The Ethical Decision Head: Operationalizing Normative Ethics in Autonomous Vehicles via Reinforcement Learning from Human Feedback

- 将伦理规则转为可微奖励,通过人类偏好学习驾驶行为
- 功利主义模型反而学会牺牲自我,与理论预期相反
- 揭示人类实际选择与道德理论存在根本差异
随着自动驾驶车辆接近SAE国际标准定义的4级和5级操作能力,其车载决策系统不仅需应对安全关键的运动控制,还需承担后续的道德权重。本文提出伦理决策头(EDH),一种基于深度强化学习的框架,将伦理推理编码为可微分奖励信号,使策略梯度代理在与CARLA仿真环境状态表示对齐的场景中学习符合道德规范的驾驶行为。评估了两种规范性伦理框架:最小化总伤亡的功利主义,以及以维持航向为绝对命令的康德主义。EDH通过近端策略优化(PPO)训练,并基于200个碰撞临界场景的人类成对偏好标注构建的布拉德利-特里奖赏模型进行学习。结果显示,在人类监督下,两种伦理框架的可学习性存在不对称性。康德条件在代码本下退化为常数预测任务,作为流水线控制,验证了训练稳定性并排除了基础设施故障对功利主义结果的解释。而功利主义代理学习到令人不安的行为:人类评分员更青睐自我牺牲而非伤亡最小化,模型也忠实学习了这一偏好。这表明,基于人类反馈的强化学习并未学习哲学定义的伦理,而是学习了人类实际生活中的伦理。
原文摘要 · Abstract (English)
As autonomous vehicles (AVs) approach Level 4 and Level 5 operational capability [SAE International, 2018], their on- board decision systems must handle not only safety-critical locomotion but also their subsequent moral weight. This paper details the Ethical Decision Head (EDH), a deep re- inforcement learning (RL) framework that encodes ethical reasoning as a differentiable reward signal, enabling a pol- icy gradient agent to learn morally-aligned driving behavior in scenarios whose state representation is aligned with the CARLA simulation environment [Dosovitskiy et al., 2017]. Two normative frameworks are instantiated and evaluated: a Utilitarian framework minimizing total casualties and a Kan- tian framework enforcing course maintenance as a categori- cal imperative. The EDH is trained via Proximal Policy Op- timization (PPO) [Schulman et al., 2017] against a Bradley- Terry reward model [Bradley and Terry, 1952] learned from pairwise human preference annotations over 200 collision- imminent scenarios. Results reveal an asymmetry in the learnability of normative ethical frameworks under human su- pervision. The Kantian condition, which reduces to a con- stant prediction task under the codebook, serves as a pipeline control: it confirms training stability and rules out infrastruc- ture failure as an explanation for the utilitarian result. The Utilitarian agent learned something more unsettling: human raters rewarded self-sacrifice over casualty minimization, and the model learned that preference faithfully. This divergence between what humans prescribe in theory and what they re- ward in practice suggests that RLHF does not learn ethics as philosophers define it, but as humans live it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。