arXiv:2410.09972cs.LGcs.AI2024-10被引 6

只重建视觉任务相关部分,让智能体在干扰中更高效学习

Make the Pertinent Salient: Task-Relevant Reconstruction for Visual Control with Distractions

  • 用分割掩码仅重构图像中的任务相关区域
  • 在带干扰的环境中样本效率提升,稀疏奖励任务也能解决
  • 适合需要抗干扰视觉控制的机器人任务

基于模型的强化学习(MBRL)在视觉控制任务中表现强劲,但面对视觉干扰时仍难以实现泛化感知。现有方法因需编码无关物体而使表征学习复杂化。本文在DREAMER基础上提出分割梦境(Segmentation Dreamer, SD),利用任务先验知识对图像施加分割掩码,仅重建任务相关区域。该策略显著降低表征学习复杂度。SD可配合真实掩码或使用有误差的分割模型,通过选择性应用重建损失避免错误信号。在修改版DeepMind Control(DMC)和Meta-World任务中加入视觉干扰后,SD在样本效率和最终性能上均优于以往方法,尤其在稀疏奖励场景下表现出色,无需繁琐奖励设计即可训练出视觉鲁棒的智能体。

原文摘要 · Abstract (English)

Recent advancements in Model-Based Reinforcement Learning (MBRL) have made it a powerful tool for visual control tasks. Despite improved data efficiency, it remains challenging to train MBRL agents with generalizable perception. Training in the presence of visual distractions is particularly difficult due to the high variation they introduce to representation learning. Building on DREAMER, a popular MBRL method, we propose a simple yet effective auxiliary task to facilitate representation learning in distracting environments. Under the assumption that task-relevant components of image observations are straightforward to identify with prior knowledge in a given task, we use a segmentation mask on image observations to only reconstruct task-relevant components. In doing so, we greatly reduce the complexity of representation learning by removing the need to encode task-irrelevant objects in the latent representation. Our method, Segmentation Dreamer (SD), can be used either with ground-truth masks easily accessible in simulation or by leveraging potentially imperfect segmentation foundation models. The latter is further improved by selectively applying the reconstruction loss to avoid providing misleading learning signals due to mask prediction errors. In modified DeepMind Control suite (DMC) and Meta-World tasks with added visual distractions, SD achieves significantly better sample efficiency and greater final performance than prior work. We find that SD is especially helpful in sparse reward tasks otherwise unsolvable by prior work, enabling the training of visually robust agents without the need for extensive reward engineering.

视觉控制强化学习抗干扰表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。