用正则化提升潜在动态预测,让行为基础模型更鲁棒。
Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models
- 在潜在空间中预测下一状态,加入正交性约束防止特征相似
- 在零样本强化学习中达到或超越现有复杂方法的性能
- 特别适合数据覆盖不足的场景,对训练数据要求更低
行为基础模型(BFM)能生成适应未知奖励或任务的智能体,但其表现受限于预设状态特征的表达能力。现有方法依赖复杂的表征学习目标和充分的数据覆盖来训练有效的特征。本文重新审视了潜在空间中的自监督下一状态预测,发现该目标易导致状态特征趋同,降低特征空间的跨度。为此提出正则化潜在动态预测(RLDP),通过引入简单正交性约束保持特征多样性,在零样本强化学习中表现与先进方法相当甚至更优。实验表明,当数据覆盖不足时,传统方法性能显著下降,而RLDP仍保持稳定表现。
原文摘要 · Abstract (English)
Behavioral Foundation Models (BFMs) produce agents with the capability to adapt to any unknown reward or task. These methods, however, are only able to produce near-optimal policies for the reward functions that are in the span of some pre-existing state features, making the choice of state features crucial to the expressivity of the BFM. As a result, BFMs are trained using a variety of complex objectives and require sufficient dataset coverage, to train task-useful spanning features. In this work, we examine the question: are these complex representation learning objectives necessary for zero-shot RL? Specifically, we revisit the objective of self-supervised next-state prediction in latent space for state feature learning, but observe that such an objective alone is prone to increasing state-feature similarity, and subsequently reducing span. We propose an approach, Regularized Latent Dynamics Prediction (RLDP), that adds a simple orthogonality regularization to maintain feature diversity and can match or surpass state-of-the-art complex representation learning methods for zero-shot RL. Furthermore, we empirically show that prior approaches perform poorly in low-coverage scenarios where RLDP still succeeds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。