提出用谱熵衡量与控制强化学习中评论家模型的复杂度。
Gauging, Measuring, and Controlling Critic Complexity in Actor-Critic Reinforcement Learning
- 用谱有效秩熵量化评论家权重矩阵的复杂度
- 复杂度变化与训练行为系统相关,且算法任务间差异明显
- 通过惩罚项可直接调控复杂度,适用于模型诊断与干预
Actor-critic方法依赖于学习到的评论家,但其质量通常仅通过回报、时序差分误差或价值损失间接评估。本文引入评论家复杂度作为新的诊断与干预维度。分析采用谱有效秩熵——一种对评论家权重矩阵奇异值分布的秩类摘要——来评估模型复杂度。在TD3和PPO实验中,同时追踪复杂度、回报与蒙特卡洛价值估计偏差。结果表明,复杂度在整个训练过程中可测量,并与训练行为系统相关,但其关系在不同算法、任务与超参数下呈现异质性。进一步通过在评论家损失中加入谱熵惩罚项,验证了复杂度可直接调控。该干预能可靠改变目标谱量,证明复杂度不仅可观测,也可控。回报变化被视为任务相关的证据,而非普遍性能提升,因整体复杂度控制效果存在差异。
原文摘要 · Abstract (English)
Actor-critic methods depend on learned critics, but critic quality is often evaluated only indirectly through return, temporal-difference error, or value loss. Critic complexity is introduced as an additional diagnostic and intervention dimension for actor-critic reinforcement learning. The analysis uses spectral effective-rank entropy, a rank-like summary of the singular-value distributions of critic weight matrices, to assess critic model complexity. Across TD3 and PPO experiments, critic complexity is tracked together with return and Monte Carlo value-estimation bias. The results show that critic complexity is measurable throughout training and is systematically associated with training behavior, while also making clear that the relationship is heterogeneous across algorithms, tasks, and hyperparameters. A direct complexity-control intervention is then evaluated by adding a spectral-entropy penalty to the critic loss. This intervention reliably changes the targeted spectral quantity, demonstrating that critic complexity can be controlled rather than only observed. Return effects are treated as task-dependent evidence rather than as a general performance claim, because overall complexity-control results vary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。