arXiv:2605.28675cs.LG2026-05

用大偏差理论优化强化学习数据采集效率,降低试错成本。

Optimal Data Acquisition for Reinforcement Learning: A Large Deviations Perspective

  • 基于大偏差理论构建数据采集的统一框架,以误差概率衰减率衡量效率。
  • 提出可计算的凸松弛方法,实现近似最优的数据自适应采集策略。
  • 适用于高成本、慢反馈场景,如医疗和商业决策系统。

在商业与医疗运营中,强化学习的数据采集面临交互成本高、速度慢且常涉及人类参与的挑战。本文提出一个统一的大偏差框架,用于无限时域强化学习中的数据采集。引入策略选择误差概率的指数衰减速率作为高效性度量,并通过马尔可夫链的大偏差理论推导出该速率的变分表征,形成嵌套优化问题。基于此表征,定义了两种互补的最优性概念。由于原问题隐式且通常不可解,本文提出显式约束的可计算凸松弛方案,并设计一种懒惰一步投影次梯度算法求解。利用迭代结果构造自适应数据采集策略,证明所提强化学习算法在最优性标准下近乎鲁棒最优,仅差一个常数因子。最后将框架扩展至线性函数逼近以提升可扩展性,数值实验验证了方法的有效性。

原文摘要 · Abstract (English)

Data acquisition efficiency is a central challenge in deploying reinforcement learning in business and healthcare operations, where interactions are costly, slow, and often involve humans in the loop. This paper develops a unified large deviations framework for data acquisition in infinite-horizon reinforcement learning. We introduce the exponential decay rate of the policy-selection error probability as a principled efficiency metric and derive a variational characterization of this rate via large deviations theory for Markov chains, yielding a nested optimization problem. Based on this characterization, we formalize two complementary notions of optimality in terms of the optimal solution of the nested problem. Because the resulting program is implicit and generally intractable, we propose a tractable convex relaxation with explicit constraints. We then develop a lazy one-step projected subgradient method to solve the relaxed problem and use its iterates to construct an adaptive data acquisition policy. We prove that the resulting reinforcement learning algorithm is near-robustly optimal under our optimality criterion, up to a constant factor. Finally, we extend the framework to linear function approximation to improve scalability, and numerical experiments support the effectiveness of the proposed approach.

强化学习数据效率大偏差自适应采集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。