通过多时间尺度学习,动态抽象环境以提升决策性能。
Performance-Driven Environment Abstraction with Multi-Timescale Learning

- 用状态聚合与共享动作分布构建可控抽象
- 压缩状态空间,提升采样效率和重规划速度
- 适合大规模强化学习中的高效决策场景
我们研究在大型马尔可夫决策过程中的性能驱动型环境抽象。不同于保留几何或拓扑结构的方法,本工作追求直接优化决策质量的抽象方式。将抽象建模为通过状态空间聚合并强制每个聚合状态内共享动作分布所获得的受控近似。针对固定划分,建立了性能保证,将值函数近似误差与动作共享带来的损失分离开来。基于此分析,提出一种多时间尺度强化学习框架,联合优化策略与树状结构的环境抽象。该算法根据Q值差异动态细化或粗化状态空间区域,在性能与抽象大小、复杂度间取得平衡。实验表明,相比演员-评论家基线,该方法实现了显著的状态压缩、更高的采样效率以及更快的重规划速度。
原文摘要 · Abstract (English)
We study performance-driven environment abstraction for decision-making in large Markov decision processes. Rather than preserving geometric or topological structure, we seek abstractions that directly optimize decision quality. We model abstraction as a controlled approximation obtained by aggregating the state space and enforcing a shared action distribution within each aggregated state. For a fixed partition, we establish a performance guarantee that separates value-function approximation error from the loss introduced by action sharing. Guided by this analysis, we develop a multi-timescale reinforcement learning framework that jointly adapts the policy and a tree-structured environment abstraction. The resulting algorithm refines and coarsens regions of the state space based on Q-value discrepancies, balancing performance against abstraction size and complexity. Empirical results demonstrate substantial state compression, improved sample efficiency, and faster replanning compared to actor-critic baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。