arXiv:2410.07441cs.LGcs.AI2024-10ICML被引 5

无需数据增强,让视觉强化学习模型零样本泛化到新环境。

Zero-Shot Generalization of Vision-Based RL Without Data Augmentation

  • 引入关联潜在解耦机制,结合联想记忆建模提升泛化能力。
  • 在复杂任务变化下实现零样本泛化,性能优于依赖数据增强的方法。
  • 为无监督泛化提供新思路,适合研究视觉强化学习的学者。

视觉强化学习(RL)代理在新环境中的泛化能力仍是重大挑战。当前主流方法依赖大规模数据集或数据增强技术以防止过拟合并提升下游泛化性能,但其计算成本和数据收集开销随任务变化数量呈指数增长,且可能加剧强化学习训练的不稳定性。本文受计算神经科学最新进展启发,提出一种名为关联潜在解耦(ALDA)的新模型,基于标准离策略强化学习框架,实现零样本泛化。具体而言,重新审视潜在变量解耦在强化学习中的作用,证明将解耦与联想记忆机制结合,可在不使用数据增强的情况下,在困难的任务变化上实现零样本泛化。此外,我们形式化证明了数据增强本质上是一种弱解耦形式,并讨论其理论意义。

原文摘要 · Abstract (English)

Generalizing vision-based reinforcement learning (RL) agents to novel environments remains a difficult and open challenge. Current trends are to collect large-scale datasets or use data augmentation techniques to prevent overfitting and improve downstream generalization. However, the computational and data collection costs increase exponentially with the number of task variations and can destabilize the already difficult task of training RL agents. In this work, we take inspiration from recent advances in computational neuroscience and propose a model, Associative Latent DisentAnglement (ALDA), that builds on standard off-policy RL towards zero-shot generalization. Specifically, we revisit the role of latent disentanglement in RL and show how combining it with a model of associative memory achieves zero-shot generalization on difficult task variations without relying on data augmentation. Finally, we formally show that data augmentation techniques are a form of weak disentanglement and discuss the implications of this insight.

视觉强化学习零样本泛化解耦表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。