arXiv:2602.18724cs.AI2026-02被引 1

用预测双生度度量提升视觉强化学习的探索效率

Task-Aware Exploration via a Predictive Bisimulation Metric

  • 通过预测奖励差值解决稀疏奖励下的表示坍塌问题
  • 在MetaWorld和Maze2D上超越最新基线,探索能力更强
  • 适合需要高效探索的复杂视觉任务场景

在稀疏奖励下加速视觉强化学习中的探索仍具挑战性,主要源于大量与任务无关的视觉变化。尽管内在探索有所进展,许多方法要么依赖低维状态,要么缺乏任务感知策略,导致在视觉领域表现脆弱。为此,我们提出TEB(Task-aware Exploration),通过预测双生度度量将任务相关表征与探索紧密耦合。具体而言,TEB不仅利用该度量学习行为基础的任务表征,还衡量学习到的潜在空间中的行为内在新颖性。为实现这一目标,我们理论上通过引入简单但有效的预测奖励差值,缓解了稀疏奖励下退化双生度度量的表示坍塌问题。基于此鲁棒度量,设计了基于势能的探索奖励,用于衡量潜在空间中相邻观测的相对新颖性。在MetaWorld和Maze2D上的大量实验表明,TEB展现出卓越的探索能力,并优于近期基线方法。

原文摘要 · Abstract (English)

Accelerating exploration in visual reinforcement learning under sparse rewards remains challenging due to the substantial task-irrelevant variations. Despite advances in intrinsic exploration, many methods either assume access to low-dimensional states or lack task-aware exploration strategies, thereby rendering them fragile in visual domains. To bridge this gap, we present TEB, a Task-aware Exploration approach that tightly couples task-relevant representations with exploration through a predictive Bisimulation metric. Specifically, TEB leverages the metric not only to learn behaviorally grounded task representations but also to measure behaviorally intrinsic novelty over the learned latent space. To realize this, we first theoretically mitigate the representation collapse of degenerate bisimulation metrics under sparse rewards by internally introducing a simple but effective predicted reward differential. Building on this robust metric, we design potential-based exploration bonuses, which measure the relative novelty of adjacent observations over the latent space. Extensive experiments on MetaWorld and Maze2D show that TEB achieves superior exploration ability and outperforms recent baselines.

强化学习探索机制视觉任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。