arXiv:2412.13662cs.CVcs.AI2024-12AAAI被引 9

对比两种视觉策略学习方法,发现状态转视觉DAgger在难任务中更稳定。

When Should We Prefer State-to-Visual DAgger Over Visual Reinforcement Learning?

  • 先学状态策略,再用在线模仿学视觉策略
  • 难任务下性能更稳定,但样本效率提升不明显
  • 适合对训练稳定性要求高的复杂视觉任务

从高维视觉输入(如像素、点云)中学习策略在诸多应用中至关重要。视觉强化学习直接从视觉观测训练策略,虽具潜力,但面临样本效率低和计算成本高的挑战。本研究在三个基准的16个任务上,对状态转视觉DAgger(两阶段框架:先训练状态策略,再通过在线模仿学习视觉策略)与视觉强化学习进行了实证比较。重点评估了两者在最终性能、样本效率和计算成本方面的表现。结果意外发现,状态转视觉DAgger并非在所有任务上都优于视觉强化学习,但在挑战性任务中表现出显著优势,性能更一致。其在样本效率上的提升较弱,但通常能减少整体训练耗时。基于此,我们为实践者提供选择建议,并期望研究成果为未来视觉策略学习研究提供有益视角。

原文摘要 · Abstract (English)

Learning policies from high-dimensional visual inputs, such as pixels and point clouds, is crucial in various applications. Visual reinforcement learning is a promising approach that directly trains policies from visual observations, although it faces challenges in sample efficiency and computational costs. This study conducts an empirical comparison of State-to-Visual DAgger, a two-stage framework that initially trains a state policy before adopting online imitation to learn a visual policy, and Visual RL across a diverse set of tasks. We evaluate both methods across 16 tasks from three benchmarks, focusing on their asymptotic performance, sample efficiency, and computational costs. Surprisingly, our findings reveal that State-to-Visual DAgger does not universally outperform Visual RL but shows significant advantages in challenging tasks, offering more consistent performance. In contrast, its benefits in sample efficiency are less pronounced, although it often reduces the overall wall-clock time required for training. Based on our findings, we provide recommendations for practitioners and hope that our results contribute valuable perspectives for future research in visual policy learning.

视觉强化学习策略学习样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。