提出软前向后向算法,让零样本强化学习能直接优化任意复杂目标。
Soft Forward-Backward Representations for Zero-shot Reinforcement Learning with General Utilities
- 用最大熵思想改进前向后向算法,从离线数据中提取随机策略族。
- 无需迭代优化,在测试时直接求解非加性奖励的复杂目标。
- 适用于分布匹配、纯探索等传统方法无法处理的任务,适合研究者参考。
零样本强化学习近年取得进展,可从无标签离线数据中提取多样行为。前向后向算法(FB)能近似解决所有加性奖励、线性于占据测度的强化学习问题。本文将此框架拓展至更一般的通用效用场景,即目标为占据测度的任意可微函数,该设定更具表达力,可涵盖分布匹配或纯探索等任务,无法简化为加性奖励。我们提出一种新的最大熵(软)前向后向算法,从离线数据中恢复出一族随机策略。结合零阶搜索对紧凑策略嵌入进行优化,该方法可避免迭代优化,直接在测试时优化通用效用。在教学性与高维实验中均验证了该方法在保持原有优势的同时,显著扩展了适用范围。
原文摘要 · Abstract (English)
Recent advancements in zero-shot reinforcement learning (RL) have facilitated the extraction of diverse behaviors from unlabeled, offline data sources. In particular, forward-backward algorithms (FB) can retrieve a family of policies that can approximately solve any standard RL problem (with additive rewards, linear in the occupancy measure), given sufficient capacity. While retaining zero-shot properties, we tackle the greater problem class of RL with general utilities, in which the objective is an arbitrary differentiable function of the occupancy measure. This setting is strictly more expressive, capturing tasks such as distribution matching or pure exploration, which may not be reduced to additive rewards. We show that this additional complexity can be captured by a novel, maximum entropy (soft) variant of the forward-backward algorithm, which recovers a family of stochastic policies from offline data. When coupled with zero-order search over compact policy embeddings, this algorithm can sidestep iterative optimization schemes, and optimizes general utilities directly at test-time. Across both didactic and high-dimensional experiments, we demonstrate that our method retains favorable properties of FB algorithms, while also extending their range to more general RL problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。