用人类神经数据校准机器人愧疚感,让合作更贴近真实人类行为。
Calibrating Artificial Guilt: Neurally Grounded Reward Shaping for Prosocial Multi-Agent Reinforcement Learning

- 从脑扫描数据中提取愧疚权重,用于强化学习奖励设计
- 校准后智能体选择安全方案的比例达0.459,接近人类的0.484
- 适合研究人类行为建模与社会性强化学习的学者
合作式多智能体强化学习常通过添加社会性奖励项来促进协作,但其权重通常人为设定。本文提出能否从人类神经与行为数据中自动校准‘愧疚’信号并迁移至机器代理。基于公开的SoDec责任任务fMRI数据集(40名参与者),我们通过固定效应回归分析瞬时幸福感变化与结果类型计数的关系,恢复出愧疚权重为伙伴负收益减去社会负收益的对比值($ ilde{w}=1.118$,Cohen's $d=0.214$)。将该权重嵌入双智能体社交彩票环境,在四种奖励塑造策略下训练独立的近端策略优化(PPO)代理:神经校准、均匀常数、零(自私)及单位系数基准。每组运行1,000次评估回合,结果显示校准组与人类安全选择率(0.484)最接近,KL散度仅0.0012;其余三组偏离程度高达一至三个数量级。表明人类神经行为先验可作为社会性奖励塑造的定量约束。
原文摘要 · Abstract (English)
Cooperative multi-agent reinforcement learning often adds social terms to individual rewards, yet the scale of those terms is usually chosen by hand. We ask whether a guilt signal can instead be calibrated from human neural and behavioural data and transferred to artificial agents. Using the public SoDec responsibility fMRI dataset (40 participants), we fit a subject-fixed-effects regression of momentary-happiness changes on outcome-type counts and recover a guilt weight as the Partner-negative minus Social-negative contrast ($\hat{w}=1.118$, Cohen's $d=0.214$). We embed this weight in a two-agent Social Lottery environment and train independent Proximal Policy Optimization actor-critics under four shaping regimes: neurally calibrated, uniform constant, zero (selfish), and a unit-coefficient oracle. Across 1{,}000 evaluation episodes per condition, the calibrated agents track the human Social safe-choice rate most closely ($0.459$ vs.\ human $0.484$; $\mathrm{KL}=0.0012$), while the other three conditions deviate by one to three orders of magnitude in KL. Human neurobehavioural priors can therefore act as quantitative constraints on prosocial reward shaping.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。