MBRL模型在部分任务上远超人类,但在另一些任务上表现极差,本文提出新评估方式与模型解决此不对称问题。
Performance Asymmetry in Model-Based Reinforcement Learning
- 按任务特性将Atari100k分为人类优势与模型优势两类子集,揭示性能不对称现象
- 最先进模型在人类优势任务上仅达人类表现的约4.8%(21倍差距),且新模型显著改善该短板
- 提出JEDI世界模型,在保持效率的同时提升人类优势任务表现,平衡整体性能
近期基于模型的强化学习(MBRL)在Atari100k基准上平均达到超人水平。然而我们发现,传统平均指标掩盖了一个关键问题:性能不对称性——某些任务中MBRL代理表现远超人类(代理最优任务),而在另一些任务上则大幅落后(人类最优任务)。尽管整体人类归一化得分(HNS)达到领先水平,但最优代理在人类最优任务上的得分却低于所有基线,人类最优与代理最优子集之间存在21倍的性能差距。为此,我们均分Atari100k为人类最优与代理最优子集,并引入更均衡的对称人类归一化得分(Sym-HNS)。进一步分析表明,当前主流像素扩散世界模型的性能不对称源于维度灾难及其在高视觉细节任务(如Breakout)中的过强表现。为此,我们提出新型潜空间端到端联合嵌入扩散(JEDI)世界模型,在Sym-HNS、人类最优任务和Breakout上均取得最佳表现,逆转了性能不对称趋势,同时提升计算效率并保持在完整Atari100k上的竞争力。
原文摘要 · Abstract (English)
Recently, Model-Based Reinforcement Learning (MBRL) have achieved super-human level performance on the Atari100k benchmark on average. However, we discover that conventional aggregates mask a major problem, Performance Asymmetry: MBRL agents dramatically outperform humans in certain tasks (Agent-Optimal tasks) while drastically underperform humans in other tasks (Human-Optimal tasks). Indeed, despite achieving SOTA in the overall mean Human-Normalized Scores (HNS), the SOTA agent scored the worst among baselines on Human-Optimal tasks, with a striking 21X performance gap between the Human-Optimal and Agent-Optimal subsets. To address this, we partition Atari100k evenly into Human-Optimal and Agent-Optimal subsets, and introduce a more balanced aggregate, Sym-HNS. Furthermore, we trace the striking Performance Asymmetry in the SOTA pixel diffusion world model to the curse of dimensionality and its prowess on high visual detail tasks (e.g. Breakout). To this end, we propose a novel latent end-to-end Joint Embedding DIffusion (JEDI) world model that achieves SOTA results in Sym-HNS, Human-Optimal tasks, and Breakout -- thus reversing the worsening Performance Asymmetry trend while improving computational efficiency and remaining competitive on the full Atari100k.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。