arXiv:2603.14600cs.LGcs.AI2026-03

通过可视化框架解析强化学习中价值估计与策略优化的互动机制。

A Loss Landscape Visualization Framework for Interpreting Reinforcement Learning: An ADHDP Case Study

  • 构建四组件可视化框架,从多角度揭示学习动态
  • 发现目标更新和训练稳定器会改变优化几何结构
  • 适合研究强化学习算法设计与训练稳定性的研究人员

强化学习算法广泛应用于动态与控制系统,但其内部学习行为的解释仍具挑战。本文在前期工作基础上,提出一个扩展的损失景观可视化框架,提供多视角学习动态分析,阐明价值估计、策略优化与时序差分(TD)信号之间的交互关系。该框架包含四个互补组件:三维重构的评论家匹配损失曲面,展示TD目标如何塑造优化几何;固定评论家下的演员损失景观,揭示策略如何利用该几何;融合时间、贝尔曼误差与策略权重的轨迹图,指示更新在曲面上的移动路径;以及状态-TD映射图,识别驱动更新的状态区域。以航天器姿态控制的行动相关启发式动态规划(ADHDP)算法为案例,应用该框架对比多个ADHDP变体,揭示训练稳定器与目标更新如何改变优化景观并影响学习稳定性。因此,该框架为跨算法设计的强化学习行为分析提供了系统且可解释的工具。

原文摘要 · Abstract (English)

Reinforcement learning algorithms have been widely used in dynamic and control systems. However, interpreting their internal learning behavior remains a challenge. In the authors' previous work, a critic match loss landscape visualization method was proposed to study critic training. This study extends that method into a framework which provides a multi-perspective view of the learning dynamics, clarifying how value estimation, policy optimization, and temporal-difference (TD) signals interact during training. The proposed framework includes four complementary components; a three-dimensional reconstruction of the critic match loss surface that shows how TD targets shape the optimization geometry; an actor loss landscape under a frozen critic that reveals how the policy exploits that geometry; a trajectory combining time, Bellman error, and policy weights that indicates how updates move across the surface; and a state-TD map that identifies the state regions that drive those updates. The Action-Dependent Heuristic Dynamic Programming (ADHDP) algorithm for spacecraft attitude control is used as a case study. The framework is applied to compare several ADHDP variants and shows how training stabilizers and target updates change the optimization landscape and affect learning stability. Therefore, the proposed framework provides a systematic and interpretable tool for analyzing reinforcement learning behavior across algorithmic designs.

强化学习损失景观可解释性ADHDP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。