用去噪模型提升视觉强化学习抗干扰能力,无需修改策略
Self-Consistent Model-based Adaptation for Visual Reinforcement Learning
- 通过去噪模型将杂乱图像转为清晰图像,增强策略鲁棒性
- 在多个基准和真实机器人数据上显著提升性能,样本效率更高
- 无监督优化去噪模型,适配各类视觉强化学习任务
视觉强化学习代理在真实应用中常因视觉干扰导致性能严重下降。现有方法依赖手工设计的增强对策略表征进行微调。本文提出自洽模型驱动适应(SCMA),一种无需修改策略的新方法。通过去噪模型将杂乱观测转换为清晰观测,SCMA可作为即插即用模块,有效缓解多种干扰。为无监督优化去噪模型,我们基于理论分析推导出最优分布匹配目标,并提出实用算法:利用预训练世界模型估计清晰观测分布以优化该目标。在多个视觉泛化基准和真实机器人数据上的大量实验表明,SCMA能有效提升各类干扰下的性能,且具备更优的样本效率。
原文摘要 · Abstract (English)
Visual reinforcement learning agents typically face serious performance declines in real-world applications caused by visual distractions. Existing methods rely on fine-tuning the policy's representations with hand-crafted augmentations. In this work, we propose Self-Consistent Model-based Adaptation (SCMA), a novel method that fosters robust adaptation without modifying the policy. By transferring cluttered observations to clean ones with a denoising model, SCMA can mitigate distractions for various policies as a plug-and-play enhancement. To optimize the denoising model in an unsupervised manner, we derive an unsupervised distribution matching objective with a theoretical analysis of its optimality. We further present a practical algorithm to optimize the objective by estimating the distribution of clean observations with a pre-trained world model. Extensive experiments on multiple visual generalization benchmarks and real robot data demonstrate that SCMA effectively boosts performance across various distractions and exhibits better sample efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。