arXiv:2602.11399cs.LGcs.AI2026-02被引 2

提出更优的无监督表示学习方法,显著提升零样本性能。

Can We Really Learn One Representation to Optimize All Rewards?

  • 基于低秩近似理论改进前向-后向表示学习
  • 误差比原方法小10万倍,零样本性能平均提升24%
  • 适合需要高效初始化的连续控制任务

随着无监督预训练在强化学习中日益普及,其理论理解的重要性与其实验成功同等重要。本文聚焦于通过交互进行无监督学习的场景,以前向-后向(FB)表示学习为代表。我们从更广泛的近期方法视角出发,将FB置于利用回归获取后续度量比低秩近似的框架中,澄清了FB表示存在的条件及低秩近似在实践中如何收敛。基于该理论,我们提出一种新变体,兼具更强理论可解释性与更易优化的特点。在教学设置及10个基于状态和图像的连续控制领域中,实验表明该方法的表示误差比原方法小10^5倍,平均零样本性能提升24%。此外,由本算法推导的零样本策略能有效作为下游任务微调的初始化方案。

原文摘要 · Abstract (English)

As unsupervised pretraining becomes increasingly ubiquitous in reinforcement learning, a more thorough theoretical understanding of these methods becomes of equal importance to their empirical success. We focus on the setting of unsupervised learning via interaction, where the forward-backward (FB) representation learning serves as a prototypical and popular example. In this paper, we shed light on FB by formally contextualizing the method within a broader class of recent methods that use regression to obtain a low-rank approximation of a successor measure ratio. Our analysis clarifies when FB representations can exist and how the low-rank approximation converges in practice. Building upon the theory, we propose a variant of FB that is both more amenable to theoretical understanding and simpler to optimize in practice. Experiments in didactic settings, as well as in $10$ state-based and image-based continuous control domains, demonstrate that our method converges to desired representations with $10^5 \times$ smaller errors than FB, achieving $+24\%$ improved zero-shot performance on average. We also demonstrate that zero-shot policies inferred by our algorithm provide an efficient initialization if the user prefers further fine-tuning on downstream tasks. Our project website is available at https://chongyi-zheng.github.io/onestep-fb.

强化学习表示学习零样本无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。