arXiv:2602.05459cs.LG2026-02

提出训练可复现性与信号可提取性评估框架,揭示高成功率不等于可靠策略提取。

Beyond Success Rates: Trainability and Extractability for Offline GCRL

  • 构建优化器学习率与AWR温度的训练可复现性图谱
  • 发现高成功率方法可能脆弱或存在低绝对天花板
  • 诊断目标区分能力与权重集中度,揭示策略生成机制

离线目标条件强化学习(GCRL)通常以最佳调参的成功率作为基准,该指标仅反映可达性能,无法揭示所学目标条件信号在政策提取中的可靠性:一种方法可能在多种值学习和提取设置下成功,也可能仅在狭窄、难寻的配置中有效。本文在共享优势加权回归(AWR)提取器下,对四种方法(GCIQL、GCIVL、QRL、CRL)进行了研究。针对每种方法,构建了随优化器学习率(影响值学习与演员优化)和AWR温度(控制演员对高优势转移的模仿选择性)变化的训练可复现性图谱。在AntMaze、Cube和Scene三个任务上观察到不同行为模式:高得分方法可能广泛可访问或极为脆弱,而宽广的相对基底仍可能处于较低绝对上限之下。通过结合事后诊断——未来目标与随机目标的区分能力及AWR权重集中度,分析其与下游成功率的关系。结果显示,在AntMaze中,未来目标与路径进展一致,这些诊断能解释图谱模式;而在Cube和Scene中,目标排序与操纵控制解耦:方法可在目标排序良好时仍失败,或在缺乏未来-随机分离的情况下通过动作条件优势成功。结果表明,峰值调优成功率不足以证明目标条件行为的广泛可提取性。训练可复现性图谱揭示此差距,而提取诊断则提供了低成本的信号转政策机制洞察。

原文摘要 · Abstract (English)

Offline goal-conditioned reinforcement learning (GCRL) is typically benchmarked by the best tuned success rate of each method. This score measures attainable performance, but it does not reveal how reliably a learned goal-conditioned signal can be extracted into a policy: a method could succeed across many value-learning and extraction settings, or only at a narrow, hard-to-find configuration. We study this gap across four methods, GCIQL, GCIVL, QRL, and CRL, under a shared advantage-weighted regression (AWR) extractor. For each method, we construct trainability landscapes over the optimizer learning rate, which affects value learning and actor optimization, and AWR temperature, which controls how selectively the actor imitates high-advantage transitions. Across AntMaze, Cube, and Scene, we observe distinct regimes: high-scoring methods may be broadly accessible or brittle, while broad relative basins may still sit below low absolute ceilings. To interpret these differences, we pair landscapes with post-hoc diagnostics of future-vs-random goal discrimination and AWR weight concentration. Their relationship to downstream success is task-dependent. On AntMaze, where future goals align with path-like progress, these diagnostics explain landscape regimes. On Cube and Scene, goal ranking and manipulation control decouple: methods can rank goals well while failing downstream, or succeed through action-conditioned advantages despite weak future-vs-random separation. These results show that peak tuned success alone does not establish broadly extractable goal-conditioned behavior. Trainability landscapes expose this gap, while extraction diagnostics offer a lower-cost lens on how learned signals become policies.

强化学习离线学习目标条件可提取性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。