提出SRCP框架,让视觉无监督强化学习更懂关键信息,零样本泛化更强。
Saliency-Guided Representation with Consistency Policy Learning for Visual Unsupervised Reinforcement Learning
- 用显著性引导动态任务,聚焦与变化相关区域,提升表示质量
- 在4个数据集16个任务上实现当前最优零样本泛化性能
- 适合研究视觉无监督强化学习、追求强泛化能力的开发者
零样本无监督强化学习(URL)为构建可泛化到未见任务的通用智能体提供了新方向。现有方法中,后继表示(SR)在结构化、低维场景中表现良好,但在高维视觉环境中难以扩展。我们通过实证分析发现:(1)SR目标常导致表示关注无关动态区域,造成后继度量不准,影响任务泛化;(2)错误表示阻碍策略建模多模态动作分布并控制技能。为此,我们提出显著性引导表示与一致性策略学习框架(SRCP),将表示学习与后继训练解耦,引入显著性引导的动力学任务以捕获与动态相关的表示,从而提升后继度量和任务泛化能力。同时,结合快速采样一致性策略、针对URL设计的无分类器引导及定制训练目标,增强条件策略建模与技能可控性。在ExORL基准的4个数据集共16个任务上的大量实验表明,SRCP在视觉URL中实现了当前最优的零样本泛化性能,且兼容多种SR方法。
原文摘要 · Abstract (English)
Zero-shot unsupervised reinforcement learning (URL) offers a promising direction for building generalist agents capable of generalizing to unseen tasks without additional supervision. Among existing approaches, successor representations (SR) have emerged as a prominent paradigm due to their effectiveness in structured, low-dimensional settings. However, SR methods struggle to scale to high-dimensional visual environments. Through empirical analysis, we identify two key limitations of SR in visual URL: (1) SR objectives often lead to suboptimal representations that attend to dynamics-irrelevant regions, resulting in inaccurate successor measures and degraded task generalization; and (2) these flawed representations hinder SR policies from modeling multi-modal skill-conditioned action distributions and ensuring skill controllability. To address these limitations, we propose Saliency-Guided Representation with Consistency Policy Learning (SRCP), a novel framework that improves zero-shot generalization of SR methods in visual URL. SRCP decouples representation learning from successor training by introducing a saliency-guided dynamics task to capture dynamics-relevant representations, thereby improving successor measure and task generalization. Moreover, it integrates a fast-sampling consistency policy with URL-specific classifier-free guidance and tailored training objectives to improve skill-conditioned policy modeling and controllability. Extensive experiments on 16 tasks across 4 datasets from the ExORL benchmark demonstrate that SRCP achieves state-of-the-art zero-shot generalization in visual URL and is compatible with various SR methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。