arXiv:2603.14589cs.LGcs.AI2026-03

将批评者匹配损失可视化扩展至离线强化学习,揭示其优化几何结构。

Adapting Critic Match Loss Landscape Visualization to Off-policy Reinforcement Learning

  • 基于固定回放缓冲区与预计算目标,适配SAC算法的批处理机制。
  • 通过锐度、盆地面积等指标量化分析收敛与发散案例的几何差异。
  • 适合研究离线强化学习优化动力学的算法开发者与控制工程师。

本工作将成熟的批评者匹配损失景观可视化方法从在线强化学习拓展至离线强化学习(RL),旨在揭示批评者学习背后的优化几何结构。离线RL与逐步在线的演员-批评者学习在数据流和目标计算上存在结构性差异。基于这两点,该方法被适配至软演员-批评者(SAC)算法,通过固定回放缓冲区及选定策略的预计算批评者目标,对损失评估进行同步调整。训练过程中记录的批评者参数投影至主成分平面,评估批评者匹配损失,形成三维景观并叠加二维优化路径。应用于航天器姿态控制问题,利用锐度、盆地面积和局部各向异性等指标,结合时间序列景观快照,对收敛与发散的SAC及发散的动作依赖启发式动态规划(ADHDP)案例进行定性与定量分析。结果表明,该适配后的可视化框架可作为基于回放缓冲的离线强化学习控制问题中批评者优化动态的几何诊断工具。

原文摘要 · Abstract (English)

This work extends an established critic match loss landscape visualization method from online to off-policy reinforcement learning (RL), aiming to reveal the optimization geometry behind critic learning. Off-policy RL differs from stepwise online actor-critic learning in its replay-based data flow and target computation. Based on these two structural differences, the critic match loss landscape visualization method is adapted to the Soft Actor-Critic (SAC) algorithm by aligning the loss evaluation with its batch-based data flow and target computation, using a fixed replay batch and precomputed critic targets from the selected policy. Critic parameters recorded during training are projected onto a principal component plane, where the critic match loss is evaluated to form a 3-D landscape with an overlaid 2-D optimization path. Applied to a spacecraft attitude control problem, the resulting landscapes are analyzed both qualitatively and quantitatively using sharpness, basin area, and local anisotropy metrics, together with temporal landscape snapshots. Comparisons between convergent SAC, divergent SAC, and divergent Action-Dependent Heuristic Dynamic Programming (ADHDP) cases reveal distinct geometric patterns and optimization behaviors under different algorithmic structures. The results demonstrate that the adapted critic match loss visualization framework serves as a geometric diagnostic tool for analyzing critic optimization dynamics in replay-based off-policy RL-based control problems.

强化学习优化几何可视化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。