长期累积损伤问题中,策略梯度易因完成度与最优性双重失效,本文提出解耦分析框架。
Completion vs Optimality: Policy Gradient in Long-Horizon Cumulative-Damage Problems

- 解耦完成度与最优性,分离两类失败模式
- 线性软惩罚下完成率下降,最优性差距达0.271
- 在积木工和篮球员两类长周期场景中验证预测
长期决策问题中,局部吸引的行为可能导致全局不利后果。本文识别出策略梯度方法在此类问题中的两种正交失败模式:完成度(达到最终时间步而非提前终止)与最优性(在完成前提下匹配动态规划基准)。在使用线性软惩罚的PPO中,仅开放时序边界会降低完成率:惩罚平衡使主导行为占比趋零;结合动作空间限制与时序访问可实现完成,但存在0.271的最优性差距,源于损伤起点处第一阶段的贪婪承诺。研究推导出四个可检验预测,并在两个独立校准环境(49步积木工生涯、20季NBA大前锋生涯)中验证,两者共享抽象结构但领域、时序、动作集与校准数据不同。所有预测定性复现。时序不变性在四组测试中三组成立,仅H=15时例外,与参数设定下的临界值区间$H^* ∈ [6,14]$一致。
原文摘要 · Abstract (English)
Long-horizon decision problems with cumulative damage couple locally attractive actions to globally adverse outcomes. We identify two orthogonal failure modes for policy-gradient methods on this class and propose a decomposition that separates them: \emph{completion} (reaching the terminal horizon rather than exiting via an implicit terminal constraint) and \emph{optimality} (matching the dynamic-programming reference given completion). Under PPO with a linear soft penalty, granting horizon access alone reduces the completion rate: the penalty's equilibrium drives the dominant-activity share to zero, while action-space restriction combined with horizon access achieves completion but leaves an optimality gap ($ΔM_{\text{final}} = 0.271$) that we trace to first-phase greedy commitment at the damage origin. We derive four testable predictions and evaluate them in two separately calibrated environments that share the same abstract structure but differ in domain, horizon, activity set, and calibration data: a 49-step bricklayer career and a 20-season NBA power-forward career. All four predictions replicate qualitatively. The horizon-invariance prediction is met at three of four tested horizons, with the exception at $H = 15$ consistent with the $H^*$ boundary ($H^* \in [6, 14]$ under the NBA parameters).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。