arXiv:2509.17092cs.LG2025-09被引 1

发现深度强化学习的难度主要由表示方式决定,传统指标失效。

On the Limits of Tabular Hardness Metrics for Deep RL: A Study with the Pharos Benchmark

  • 用新工具Pharos系统控制环境与表示方式,对比不同输入下的学习难度
  • 相同任务下,像素输入比状态向量难得多,表征是关键影响因素
  • 提出需发展考虑表示能力的新评估标准,适合研究基准设计的学者

深度强化学习(RL)的评估体系尚不完善,远落后于表格型RL的理论驱动基准。尽管表格型设置有明确的难度度量如MDP直径和次优差距,但深度RL基准多凭直觉和流行选择。本文探究这些表格型度量能否用于非表格型场景,揭示根本性差距:非表格环境的难度主要受表征硬度支配。同一潜在MDP,当智能体接收状态向量或像素观测时,挑战程度天差地别。为此,我们引入开源库Pharos,支持对环境结构与智能体表示进行系统控制。基于Pharos的广泛案例研究显示,尽管表格度量提供部分参考,但无法独立准确预测深度RL智能体性能。该工作强调亟需发展考虑表征的新硬度度量,并将Pharos定位为构建此类度量的关键工具。

原文摘要 · Abstract (English)

Principled evaluation is critical for progress in deep reinforcement learning (RL), yet it lags behind the theory-driven benchmarks of tabular RL. While tabular settings benefit from well-understood hardness measures like MDP diameter and suboptimality gaps, deep RL benchmarks are often chosen based on intuition and popularity. This raises a critical question: can tabular hardness metrics be adapted to guide non-tabular benchmarking? We investigate this question and reveal a fundamental gap. Our primary contribution is demonstrating that the difficulty of non-tabular environments is dominated by a factor that tabular metrics ignore: representation hardness. The same underlying MDP can pose vastly different challenges depending on whether the agent receives state vectors or pixel-based observations. To enable this analysis, we introduce \texttt{pharos}, a new open-source library for principled RL benchmarking that allows for systematic control over both environment structure and agent representations. Our extensive case study using \texttt{pharos} shows that while tabular metrics offer some insight, they are poor predictors of deep RL agent performance on their own. This work highlights the urgent need for new, representation-aware hardness measures and positions \texttt{pharos} as a key tool for developing them.

强化学习基准测试表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。