arXiv:2607.04153cs.LGcs.AI2026-07

用遮蔽预测提升视觉强化学习的样本效率

Mask-based Predictive Representations for Reinforcement Learning

论文配图:Mask-based Predictive Representations for Reinforcement Learning
图 1 · 摘自论文原文
  • 通过遮蔽图像序列中的信息,让模型预测被遮部分,学习更有效的状态表示
  • 在多个连续与离散控制任务上,样本效率优于当前最优方法
  • 适合研究高效强化学习、视觉表征学习的科研人员

基于视觉的深度强化学习面临高维图像输入和有限样本的挑战,亟需从图像中抽象出有效状态以实现高效的样本利用。受自然语言处理和计算机视觉启发,我们提出一种基于遮蔽预测的自监督辅助任务。该非重建方法利用智能体从环境中收集的序列信息及上下文,预测被遮蔽内容,从而增强对任务的理解并学习有效表征。结合Transformer架构,模型在隐空间中重构被遮蔽的输入序列。将该方法学习到的压缩表示用于强化学习模型后,显著提升了样本效率。在多个连续与离散控制基准测试中,该方法超越了当前最先进的样本高效强化学习方法。

原文摘要 · Abstract (English)

Vision-based deep reinforcement learning involves dealing with high-dimensional inputs of image information. It is crucial to abstract effective states from high-dimensional image inputs and limited samples for sample-efficient reinforcement learning. To address this challenge, inspired by fields such as natural language processing and computer vision, we propose a self-supervised task based on mask prediction as an auxiliary task for reinforcement learning. This non-reconstruction method uses the sequence information collected by the agent from the environment and the context information in the sequence to predict the masked information, thereby strengthening the agent's understanding of the task and learning effective representations. Combined with transformers, we find that the model reconstructs the masked input sequence in the latent space. By feeding the compressed representations learned by this method into reinforcement learning models, we observe an improvement in the sample efficiency of reinforcement learning. Moreover, the model outperforms state-of-the-art sample-efficient reinforcement learning methods on multiple continuous and discrete control benchmarks.

强化学习视觉表征自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。