arXiv:2510.04088cs.LGcs.AI2025-10被引 20

从历史数据中学习最优策略,无需环境交互

Offline Reinforcement Learning in Large State Spaces: Algorithms and Guarantees

  • 基于函数逼近的表达性假设设计算法
  • 在不同数据覆盖条件下实现学习保证
  • 适合研究离线强化学习理论与算法者

本文建立了大规模状态空间下离线强化学习的理论框架,通过历史数据学习最优策略,无需与环境在线交互。核心概念包括函数逼近的表达性假设(如贝尔曼完备性与可实现性)和数据覆盖条件(如全策略覆盖与单策略覆盖)。文中系统描述了多种算法及其性能保证,其效果取决于所作假设及对样本复杂度与计算复杂度的要求。同时探讨了开放问题及与其他相关领域的联系。

原文摘要 · Abstract (English)

This article introduces the theory of offline reinforcement learning in large state spaces, where good policies are learned from historical data without online interactions with the environment. Key concepts introduced include expressivity assumptions on function approximation (e.g., Bellman completeness vs. realizability) and data coverage (e.g., all-policy vs. single-policy coverage). A rich landscape of algorithms and results is described, depending on the assumptions one is willing to make and the sample and computational complexity guarantees one wishes to achieve. We also discuss open questions and connections to adjacent areas.

离线RL强化学习理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。