研究多智能体强化学习中奖励分配方式对模型表征的影响。
Feedback Attribution and Representation Geometry: Metrics for Comparing Individual and Shared Rewards in MARL

- 提出有效秩和动作分布差异两个轻量诊断指标。
- 观察信息主导表征几何,奖励分配影响主要体现在行为上。
- 适合关注多智能体协作机制的算法研究者阅读。
合作式多智能体强化学习系统常采用团队平均奖励,即每个智能体均获得团队整体结果,而不区分其个体贡献。本文探讨这种奖励分配方式是否在学习到的表征中留下可测量的几何或行为痕迹。提出EffRank/n(归一化有效秩)和D_act(智能体间动作分布的平均成对KL散度)作为低开销诊断指标,并在SMACv2的protoss_5_vs_5场景中测试了具备能力的MAPPO智能体。实验对比了观察单位类型与否、个体伤害贡献奖励与共享团队奖励两种设置:当单位类型被观测时,共享与个体奖励下的EffRank/n分别为0.31±0.03和0.29±0.02,探针准确率分别为0.75±0.05和0.73±0.05(显著高于1/3随机水平),而D_act在个体奖励下更高(1.23±0.06 vs. 1.07±0.20)。当单位类型被遮蔽时,探针信号下降至0.49。结论:个体奖励使智能体更具角色可分性,但表征几何主要由观察信息决定,奖励分配的影响更多体现于行为层面。因此几何诊断需控制已知角色信息,并检验未直接观测的持续角色。这两个指标增加的计算开销小于5%。
原文摘要 · Abstract (English)
Cooperative multi-agent RL systems routinely use team-averaged rewards, a feedback-attribution choice that gives each agent the team outcome regardless of its individual contribution. We ask whether this leaves a measurable signature, geometric or behavioral, on learned representations. We propose EffRank/$n$ (effective rank normalized by agent count) and $D_\text{act}$ (mean pairwise KL divergence between agents' action distributions) as low-overhead diagnostics for reward-attribution effects, then test them on competent MAPPO agents in SMACv2 \texttt{protoss\_5\_vs\_5}, where unit type is encoded in the observation. In an observation $\times$ reward-attribution comparison (unit type observed vs.\ masked; individual damage-contribution reward vs.\ shared team reward), geometry follows observation rather than reward. With unit type observed, shared and individual rewards have similar EffRank/$n$ ($0.31{\pm}0.03$ vs.\ $0.29{\pm}0.02$) and probe accuracy ($0.75{\pm}0.05$ vs.\ $0.73{\pm}0.05$, both $\gg 1/3$ chance), while $D_\text{act}$ leans higher under individual rewards ($1.23{\pm}0.06$ vs.\ $1.07{\pm}0.20$). Masking unit type cuts the above-chance probe signal by more than half, to $0.49$ in both reward arms. In short: individually rewarded agents are competent and separable by role, but on SMACv2 the observation explains the geometry and reward attribution shows up mainly in behavior. Thus geometric diagnostics must control for observed role information and test persistent roles that are not directly observed. EffRank/$n$ and $D_\text{act}$ add $<$5\% overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。