提出新方法让机器人在离线强化学习中更精准选择动作。
Dual Advantage Fields

- 用双目标表示构建全局可达性价值场,再提取局部动作优势信号。
- 在多种任务上提升整体可信赖指标,尤其在非直接路径场景表现优。
- 适合离线强化学习中的复杂导航与操作任务,如机器人控制。
离线目标条件强化学习需要长期可达性估计和局部动作比较。双目标表示能捕捉全局目标可达性,但无法直接给出特定状态下的优选动作。本文提出双优势场(Dual Advantage Fields),将双线性双值模型转化为局部优势信号。在双线性双参数化下,目标嵌入是价值场对状态表示的梯度。DAF学习一个动作效应模型,预测动作带来的折扣特征位移,并通过该位移与目标方向的对齐程度评分。在可实现情况下,该评分等于目标条件贝尔曼优势,提供标准局部策略改进保证。在OGBench的运动、操作和谜题任务上,DAF提升了综合可信赖指标,在局部最优动作与直接朝向目标不同的场景中表现优异。
原文摘要 · Abstract (English)
Offline goal-conditioned reinforcement learning requires both long-horizon reachability estimates and local action comparisons. Dual goal representations provide value fields that capture global goal reachability, but they do not directly specify which action should be preferred at a given state. We propose Dual Advantage Fields, a policy-extraction method that turns a bilinear dual value model into a local advantage signal. Under bilinear dual parameterization, the goal embedding is the gradient of the value field with respect to the state representation. DAF learns an action-effect model that predicts the discounted feature displacement induced by an action and scores actions by the alignment between this displacement and the goal direction. In the realizable case, this score equals the goal-conditioned Bellman advantage, yielding a standard local policy-improvement guarantee. On OGBench locomotion, manipulation, and puzzle tasks, DAF improves aggregate RLiable metrics and performs strongly in settings where locally correct actions differ from direct movement toward the final goal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。