轨迹价值受算法影响:同一轨迹在不同算法中价值相反。
Algorithm-Relative Trajectory Valuation in Policy Gradient Control
- 用轨迹谢泼德值分析强化学习中轨迹贡献
- 原始REINFORCE下,激励持续性越高价值越低(r≈-0.38)
- 算法稳定后相关性反转,适合算法选型与数据清洗
我们研究了策略梯度控制中轨迹价值如何依赖于学习算法。在不确定的LQR环境中使用轨迹谢泼德值,发现原始REINFORCE下,激励持续性(PE)与边际价值呈负相关(r≈-0.38)。我们证明了方差介导机制:(i) 固定能量下,更高PE降低梯度方差;(ii) 接近鞍点时,更高方差增加逃逸概率,提升边际贡献。当系统通过状态白化或费舍尔预条件稳定后,该方差通道被中和,信息量主导,相关性转为正(r≈+0.29)。因此,轨迹价值具有算法相对性。实验验证了该机制,并表明留一法评分与谢泼德值互补用于剪枝,而谢泼德值可识别有毒子集。
原文摘要 · Abstract (English)
We study how trajectory value depends on the learning algorithm in policy-gradient control. Using Trajectory Shapley in an uncertain LQR, we find a negative correlation between Persistence of Excitation (PE) and marginal value under vanilla REINFORCE ($r\approx-0.38$). We prove a variance-mediated mechanism: (i) for fixed energy, higher PE yields lower gradient variance; (ii) near saddles, higher variance increases escape probability, raising marginal contribution. When stabilized (state whitening or Fisher preconditioning), this variance channel is neutralized and information content dominates, flipping the correlation positive ($r\approx+0.29$). Hence, trajectory value is algorithm-relative. Experiments validate the mechanism and show decision-aligned scores (Leave-One-Out) complement Shapley for pruning, while Shapley identifies toxic subsets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。