arXiv:2502.05555cs.CV2025-02AAAI被引 4

自适应预训练视觉编码器提升视觉强化学习的泛化与采样效率

Efficient Reinforcement Learning Through Adaptively Pretrained Visual Encoder

  • 通过自适应数据增强策略优化编码器预训练
  • 仅需少量环境交互即可达到主流RL模型性能
  • 适合追求高效视觉强化学习的科研与工程应用

尽管强化学习(RL)代理能成功处理复杂任务,但将已学技能泛化到陌生环境仍具挑战。原因之一是所用视觉编码器依赖特定任务,难以在不同场景中有效提取特征。现有研究尝试用多样化视觉输入预训练编码器以提升性能,但通常沿用现成预训练模型,未深入探索预训练周期的影响。本文提出APE框架:利用预训练阶段的自适应增强策略,在策略学习阶段仅需少量环境交互即可提取可泛化的特征。实验在DeepMind Control Suite、Atari Games和Memory Maze等多个基准上验证有效性。结果表明,主流RL方法如DreamerV3和DrQ-v2在引入APE后达到当前最优表现。此外,仅使用视觉输入时,APE显著提升采样效率,部分控制任务接近基于状态的方法表现。这表明自适应编码器预训练对提升视觉强化学习的泛化能力与效率具有巨大潜力。

原文摘要 · Abstract (English)

While Reinforcement Learning (RL) agents can successfully learn to handle complex tasks, effectively generalizing acquired skills to unfamiliar settings remains a challenge. One of the reasons behind this is the visual encoders used are task-dependent, preventing effective feature extraction in different settings. To address this issue, recent studies have tried to pretrain encoders with diverse visual inputs in order to improve their performance. However, they rely on existing pretrained encoders without further exploring the impact of pretraining period. In this work, we propose APE: efficient reinforcement learning through Adaptively Pretrained visual Encoder -- a framework that utilizes adaptive augmentation strategy during the pretraining phase and extracts generalizable features with only a few interactions within the task environments in the policy learning period. Experiments are conducted across various domains, including DeepMind Control Suite, Atari Games and Memory Maze benchmarks, to verify the effectiveness of our method. Results show that mainstream RL methods, such as DreamerV3 and DrQ-v2, achieve state-of-the-art performance when equipped with APE. In addition, APE significantly improves the sampling efficiency using only visual inputs during learning, approaching the efficiency of state-based method in several control tasks. These findings demonstrate the potential of adaptive pretraining of encoder in enhancing the generalization ability and efficiency of visual RL algorithms.

强化学习视觉编码器自适应预训练采样效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。