揭示目标条件化生成模型的不连贯性,提出通过在线微调改善策略性能。
Incoherence in Goal-Conditioned Autoregressive Models
- 通过在线重训练降低策略不连贯性,提升回报表现。
- 证明重训练过程可使策略轨迹趋于稳定,提升决策一致性。
- 适合关注强化学习生成模型与策略优化的研究者阅读。
我们从数学上研究了不连贯性:一种在朴素目标条件化自回归模型中产生的结构问题。聚焦于对自身动作进行再训练的过程,即使用在线强化学习对离线学习的策略进行微调。我们证明该过程能减少不连贯性并提升回报,并试图刻画由此产生的策略演化轨迹。通过重新定义控制即推理与软Q学习的标准概念,我们建立了三者之间的对应关系:将后验折叠进奖励函数,以及在确定性情形下降低温度参数;该对应关系具有计算意义,体现在训练-推理权衡中。通过软条件化生成模型,我们探讨了不连贯性与有效时域之间的联系。
原文摘要 · Abstract (English)
We investigate mathematically the notion of incoherence: a structural issue with reinforcement learning policies derived by naive goal-conditioning of autoregressive models. We focus on the process of re-training models on their own actions, that is, fine-tuning offline-learned policies with online RL. We prove that it decreases incoherence and leads to an improvement in return, and we aim to characterize the resulting trajectory of policies. By re-framing standard notions of control-as-inference and soft Q learning, we establish a three-way correspondence with two other ways of understanding the iterative re-training process: as folding the posterior into the reward and, in the deterministic case, as decreasing the temperature parameter; the correspondence has computational content via the training-inference trade-off. Through soft-conditioning generative models, we discuss the link between incoherence and the effective horizon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。