arXiv:2606.18531stat.MLcs.LG2026-06

研究轨迹级监督下离线强化学习何时高效,揭示其统计瓶颈与可行条件。

When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning?

  • 提出OPAC算法,从轨迹标签学习隐式奖励并优化策略。
  • 理论证明样本复杂度为$\widetilde O(H^2\sqrt{C_{sa}(π^\star)/n})$,匹配下界。
  • 发现非线性聚合目标不可学习,但结构系数控制下可实现多项式样本效率。

离线强化学习通常在过程级奖励监督下分析,但许多序列决策数据集仅记录轨迹级结果。本文建立从此类结果级监督进行离线策略优化的统计理论。首先研究经典设定:目标仍是期望累计奖励,但每条离线轨迹仅提供一个标量标签,其条件均值为累计回报。提出OPAC算法,通过学习隐式奖励模型并基于轨迹级标签优化策略。证明高概率保证为$\widetilde O(H^2\sqrt{C_{sa}(π^\star)/n})$,并给出匹配下界,刻画用轨迹级标签替代过程级奖励的精确统计代价。进一步将原理扩展至偏好反馈,保持主导的时域和集中性依赖关系,仅受偏好模型常数影响。最后研究广义结果级离线强化学习,其中监督与目标均为由潜在每步奖励非线性聚合诱导的轨迹级量。该问题一般不可学习:对全成功目标,任何离线学习器即使在确定性转移和常数集中性下仍需$Ω(2^H)$条轨迹。随后通过两个结构系数$κ_μ(σ)$和$χ_μ(σ)$识别出可解区域,捕捉结果聚合与广义贝尔曼更新中的信息损失,此时广义OPAC实现多项式样本复杂度。综上,本工作明确划分了结果级监督能否实现高效离线控制的边界。

原文摘要 · Abstract (English)

Offline reinforcement learning is typically analyzed under process-level reward supervision, yet many sequential decision datasets record only trajectory-level outcomes. We develop a statistical theory for offline policy optimization from such outcome-level supervision. We first study the canonical setting where the target remains the expected cumulative reward, but each offline trajectory provides only a scalar label whose conditional mean is the cumulative return. We propose OPAC, a pessimistic actor-critic algorithm that learns a latent reward model and optimizes a policy from trajectory-level labels. We prove a high-probability guarantee of order $\widetilde O(H^2\sqrt{C_{sa}(π^\star)/n})$ and a matching lower bound, characterizing the sharp statistical cost of replacing process-level rewards with one trajectory-level label. We then extend the principle to preference-based feedback, preserving the leading horizon and concentrability dependence up to preference-model constants. Finally, we study generalized outcome-based offline RL, where both the supervision and the objective are trajectory-level quantities induced by a nonlinear aggregation of latent per-step rewards. This problem is not learnable in general: for all-success objectives, any offline learner may require $Ω(2^H)$ trajectories even with deterministic transitions and constant concentrability. We then identify a tractable regime through two structural coefficients, $κ_μ(σ)$ and $χ_μ(σ)$, capturing information loss in outcome aggregation and generalized Bellman updates, under which generalized OPAC achieves polynomial sample complexity. Together, our results delineate when outcome-level supervision enables sample-efficient offline control and when missing process-level rewards create fundamental statistical barriers.

强化学习离线学习统计理论轨迹监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。