arXiv:2509.04970cs.ROcs.AI2025-09

用深度图引导视觉强化学习,提升机器人操作的泛化与可解释性

DeGuV: Depth-Guided Visual Reinforcement Learning for Generalization and Interpretability in Manipulation

  • 基于深度图生成可学习掩码,保留关键视觉信息
  • 在数据增强下仍保持高效训练,零样本仿真到现实迁移成功
  • 能可视化关注区域,适合需要可解释性的机器人任务

强化学习(RL)代理可从视觉输入中学习复杂任务,但在新环境中的泛化仍是重大挑战,尤其在机器人领域。尽管数据增强有助于泛化,但常牺牲样本效率和训练稳定性。本文提出DeGuV框架,通过可学习的掩码网络从深度输入生成掩码,仅保留关键视觉信息,丢弃无关像素,使代理聚焦于重要特征,增强数据增强下的鲁棒性。同时引入对比学习并稳定增强下的Q值估计,进一步提升样本效率与训练稳定性。我们在使用Franka Emika机器人的RL-ViGen基准上评估该方法,实现零样本仿真到现实迁移。结果表明,DeGuV在泛化和样本效率上均优于现有方法,并通过突出视觉输入中最相关区域提升了可解释性。

原文摘要 · Abstract (English)

Reinforcement learning (RL) agents can learn to solve complex tasks from visual inputs, but generalizing these learned skills to new environments remains a major challenge in RL application, especially robotics. While data augmentation can improve generalization, it often compromises sample efficiency and training stability. This paper introduces DeGuV, an RL framework that enhances both generalization and sample efficiency. In specific, we leverage a learnable masker network that produces a mask from the depth input, preserving only critical visual information while discarding irrelevant pixels. Through this, we ensure that our RL agents focus on essential features, improving robustness under data augmentation. In addition, we incorporate contrastive learning and stabilize Q-value estimation under augmentation to further enhance sample efficiency and training stability. We evaluate our proposed method on the RL-ViGen benchmark using the Franka Emika robot and demonstrate its effectiveness in zero-shot sim-to-real transfer. Our results show that DeGuV outperforms state-of-the-art methods in both generalization and sample efficiency while also improving interpretability by highlighting the most relevant regions in the visual input

强化学习机器人操作可解释性深度感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。