检验分布强化学习的风控声称是否真实,发现多数声称是训练幻觉。
Auditing the Risk Claims of Distributional Reinforcement Learning

- 用最优动作间Wasserstein差距检测风险宣称的真实性
- 33次实验中95%的最强风险主张被推翻,接近随机水平
- 适用于关注强化学习安全性和可信度的研究者
分布强化学习代理学习完整的回报分布,并被用于可解释性、风险敏感控制和安全监控。我们提出一个理论预期但未直接测量的问题:训练后分布代理的风险声称是否真实?审计结合决策相关筛选指标(前两个动作间的超额Wasserstein差距,等于一阶随机占优被违反的质量)、来自快照重启蒙特卡洛的真值,以及统计工具(置换零假设、自助重审、错误发现率控制),否则审计自身会制造虚假结论。在MinAtar上对QR-DQN、C51和IQN进行33次运行测试,95%置信水平下40%-95%的最强风险权衡声称被驳回;最强声称的位置与盲视真值无统计差异,几乎无任何声称可确认。这些代理所学“风险”反映的是训练伪影而非环境随机性。该伪影为结构性的(训练早期即形成,与最终得分无关,每种子独有),在完整Atari规模下仍存在,所有预训练近前沿QR-DQN在Breakout中的顶级风险声称均被驳回。已知幅度的正向对照验证了96%-100%真实声称的识别能力(相关性0.89-0.92):读取机制准确衡量代理表现,而非审计本身。依据头部的CVaR建议在最显著状态行动,效果从有益到显著劣于随机。无论是否训练风险感知或集成,伪影均无法消除;校准仅通过消除声称通过审计,表明头部无信息量,而非仅失准。我们发布审计工具包,并揭示两个曾导致误导性但看似合理的审计误判陷阱。
原文摘要 · Abstract (English)
Distributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability, risk-sensitive control, and safety monitoring. We ask a question theory anticipates but that has not been measured directly: are the risk claims of a trained distributional agent true? Our audit combines a decision-relevant screening metric (the excess Wasserstein gap between the top two actions, which equals the mass by which first-order stochastic dominance is violated), ground truth from snapshot-restart Monte Carlo, and a statistical harness (permutation nulls, bootstrap refutation, FDR control) without which the audit itself manufactures false conclusions. Across QR-DQN, C51, and IQN on MinAtar (33 runs), 40-95% of the strongest claimed risk trade-offs are refuted at 95% confidence, the placement of the strongest claims is statistically indistinguishable from truth-blind, and essentially no claim is confirmable: for these agents, the learned "risk" reflects a training artifact rather than environment stochasticity. The artifact is structural (fully formed early in training, uncorrelated with final score, idiosyncratic to each seed) and appears unchanged at full-Atari scale, with every top Breakout claim of a pretrained near-state-of-the-art QR-DQN refuted. Positive controls of known magnitude confirm 96-100% of real claims (correlation 0.89-0.92): the reading measures the agents, not the audit. Acting on the heads' CVaR advice at their most-flagged states ranges from beneficial to significantly worse than chance. Neither training for risk nor ensembling removes the artifact, and recalibration passes the audit only by nullifying the claims: the head is uninformative, not merely miscalibrated. We release the toolkit and document two silent pitfalls that produced convincing but wrong audits of our own.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。