arXiv:2410.14038cs.LG2024-10ICML被引 8

新基准SPGym让视觉强化学习的表征能力可量化评估

Sliding Puzzles Gym: A Scalable Benchmark for State Representation in Visual Reinforcement Learning

  • 用可调网格和图像池控制视觉表征难度
  • 图像多样性增加时所有算法性能下降
  • 简单数据增强反而常优于复杂表征方法

有效的视觉表征学习对强化学习(RL)智能体从原始感知输入中提取任务相关信息并跨环境泛化至关重要。然而,现有RL基准无法将表征学习能力与其他学习挑战分离评估。为此,我们提出滑动拼图环境(Sliding Puzzles Gym, SPGym),将经典的8-拼图问题转化为一个视觉强化学习任务,使用任意大规模数据集中的图像作为观测。SPGym的核心创新在于通过调节网格尺寸和图像池大小,精确控制表征学习复杂度,同时保持环境动态、观测空间和动作空间不变。这一设计使研究者能独立地隔离并扩展视觉表征挑战。通过对无模型和基于模型的多种算法进行广泛实验,我们发现当前方法在应对视觉多样性方面存在根本局限:随着图像池规模增大,所有算法均出现分布内与分布外性能退化,且复杂的表征学习技术常不如简单的数据增强策略。这些结果揭示了视觉表征学习在强化学习中的关键短板,并确立了SPGym作为推动鲁棒、可泛化决策系统发展的有力工具。

原文摘要 · Abstract (English)

Effective visual representation learning is crucial for reinforcement learning (RL) agents to extract task-relevant information from raw sensory inputs and generalize across diverse environments. However, existing RL benchmarks lack the ability to systematically evaluate representation learning capabilities in isolation from other learning challenges. To address this gap, we introduce the Sliding Puzzles Gym (SPGym), a novel benchmark that transforms the classic 8-tile puzzle into a visual RL task with images drawn from arbitrarily large datasets. SPGym's key innovation lies in its ability to precisely control representation learning complexity through adjustable grid sizes and image pools, while maintaining fixed environment dynamics, observation, and action spaces. This design enables researchers to isolate and scale the visual representation challenge independently of other learning components. Through extensive experiments with model-free and model-based RL algorithms, we uncover fundamental limitations in current methods' ability to handle visual diversity. As we increase the pool of possible images, all algorithms exhibit in- and out-of-distribution performance degradation, with sophisticated representation learning techniques often underperforming simpler approaches like data augmentation. These findings highlight critical gaps in visual representation learning for RL and establish SPGym as a valuable tool for driving progress in robust, generalizable decision-making systems.

强化学习表征学习基准测试视觉智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。