研究如何让强化学习模型忽略干扰,提升鲁棒性。
Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments
- 用统一框架评估五种度量学习方法,对比其在不同设计下的表现。
- 在370种任务配置中验证,度量学习可显著提升模型抗噪能力。
- 提出独立评估设置,揭示度量学习对表征的真实贡献,适合相关研究者。
状态抽象的关键方法是近似观测空间中的行为度量(尤其是双模拟度量),并将这些学习到的距离嵌入表示空间。尽管该方法对任务无关噪声具有鲁棒性,但准确估计这些度量仍具挑战性,需权衡多种设计选择,导致理论与实践之间存在差距。以往评估主要关注最终回报,未能阐明所学度量的质量及性能提升的来源。为系统评估度量学习在深度强化学习中的作用,我们评估了五种近期方法,它们在概念上统一为具有不同设计选择的等距嵌入。我们在20个基于状态和14个基于像素的任务上进行基准测试,涵盖370种任务配置及多种噪声设置。除最终回报外,引入去噪因子以量化编码器过滤干扰的能力。为进一步隔离度量学习的影响,我们提出并评估了一个独立度量估计设置,其中编码器仅受度量损失影响。最后,我们开源了一个模块化代码库,以提高可复现性,并支持未来在深度强化学习中度量学习的研究。
原文摘要 · Abstract (English)
A key approach to state abstraction is approximating behavioral metrics (notably, bisimulation metrics) in the observation space and embedding these learned distances in the representation space. While promising for robustness to task-irrelevant noise, as shown in prior work, accurately estimating these metrics remains challenging, requiring various design choices that create gaps between theory and practice. Prior evaluations focus mainly on final returns, leaving the quality of learned metrics and the source of performance gains unclear. To systematically assess how metric learning works in deep reinforcement learning (RL), we evaluate five recent approaches, unified conceptually as isometric embeddings with varying design choices. We benchmark them with baselines across 20 state-based and 14 pixel-based tasks, spanning 370 task configurations with diverse noise settings. Beyond final returns, we introduce the evaluation of a denoising factor to quantify the encoder's ability to filter distractions. To further isolate the effect of metric learning, we propose and evaluate an isolated metric estimation setting, in which the encoder is influenced solely by the metric loss. Finally, we release an open-source, modular codebase to improve reproducibility and support future research on metric learning in deep RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。