arXiv:2504.04164cs.LG2025-04

解决视觉强化学习中的信息冲突,提升抗干扰能力

MInCo: Mitigating Information Conflicts in Distracted Visual Model-based Reinforcement Learning

  • 用无负样本对比学习减少视觉表示冲突
  • 在动态背景干扰下实现更鲁棒的策略性能
  • 适合需要抗干扰的机器人控制任务

现有视觉模型基于强化学习(MBRL)方法在观测重建时常因信息冲突导致难以学习紧凑表示,尤其在存在无关视觉干扰时表现更差。本文从信息论角度揭示,当前方法的信息冲突源于视觉表示学习与潜在动态建模之间的不一致。为此提出新算法MInCo,通过无负样本对比学习缓解信息冲突,帮助学习对背景噪声不变的表示和鲁棒策略。为避免视觉表示学习主导训练过程,引入随时间变化的重加权机制,逐步引导学习偏向动态建模。在多个带动态背景干扰的机器人控制任务上评估,实验表明MInCo能有效学习背景噪声不变的表示,并持续优于现有最先进视觉MBRL方法。代码已公开于https://github.com/ShiguangSun/minco。

原文摘要 · Abstract (English)

Existing visual model-based reinforcement learning (MBRL) algorithms with observation reconstruction often suffer from information conflicts, making it difficult to learn compact representations and hence result in less robust policies, especially in the presence of task-irrelevant visual distractions. In this paper, we first reveal that the information conflicts in current visual MBRL algorithms stem from visual representation learning and latent dynamics modeling with an information-theoretic perspective. Based on this finding, we present a new algorithm to resolve information conflicts for visual MBRL, named MInCo, which mitigates information conflicts by leveraging negative-free contrastive learning, aiding in learning invariant representation and robust policies despite noisy observations. To prevent the dominance of visual representation learning, we introduce time-varying reweighting to bias the learning towards dynamics modeling as training proceeds. We evaluate our method on several robotic control tasks with dynamic background distractions. Our experiments demonstrate that MInCo learns invariant representations against background noise and consistently outperforms current state-of-the-art visual MBRL methods. Code is available at https://github.com/ShiguangSun/minco.

强化学习视觉建模抗干扰机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。