用敏感度分析揭示强化学习模型内部演化轨迹
Interpreting Reinforcement Learning Agents with Susceptibilities

- 将敏感度方法扩展至强化学习的损失函数扰动分析
- 在网格世界中发现策略演进中不可见的参数空间特征
- 适合研究模型内在机制与对齐训练的学者
敏感度是一种神经网络可解释性技术,通过研究可观测量后验期望值对损失扰动的响应来分析模型。本文将该方法推广至深度强化学习中的后悔(regret)场景,并在具有复杂分阶段发展的简单网格世界模型中验证其有效性。结果表明,敏感度能揭示仅通过策略演化无法捕捉的模型在参数空间中的内部特征。通过激活引导(activation-steering)验证了这些发现,并讨论了该框架在强化学习人类反馈(RLHF)后训练阶段的拓展潜力。
原文摘要 · Abstract (English)
Susceptibilities are a technique for neural network interpretability that studies the response of posterior expectation values of observables to perturbations of the loss. We generalize this construction to the setting of the regret in deep reinforcement learning and investigate the utility of susceptibilities in a simple gridworld model that nevertheless exhibits non-trivial stagewise development. We argue that susceptibilities reveal internal features of the development of the model in parameter space that one cannot detect purely by studying the development of the learned policy. We validate these results with activation-steering, and discuss the framework's extension to RLHF post-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。