arXiv:2605.03393stat.MLcs.LG2026-05被引 1

提出首个无需平稳性假设的离线上下文MDP自适应估计算法

Adaptive Estimation and Optimal Control in Offline Contextual MDPs without Stationarity

  • 基于T-估计技术构建自适应估计器,处理非平稳与模型不规则问题
  • 在两种损失函数下获得最优风险界,实现有限样本下的成本优化
  • 适用于医疗决策、推荐系统等非平稳场景的离线强化学习

上下文马尔可夫决策过程(Contextual MDPs)在生物统计学和机器学习等领域具有广泛应用。然而,将其应用于离线数据集时面临挑战,因缺乏稳健且具备理论保障的方法。本文提出一种新型自适应估计与成本优化方法,据我们所知,这是首个无需平稳性假设的此类方法,具备强最优性保证。通过克服上下文MDP固有的非平稳性与模型不规则性等关键技术难题,利用近期强大的统计技术T-估计(Baraud, 2011),在完全一般条件下建立理论保证。首先,给出基于样本选择估计器的程序,并在两种有意义的损失函数下推导出关于最优风险的界。随后,借助前述密度估计解决最优控制问题,并提供成本函数的有限样本保证。

原文摘要 · Abstract (English)

Contextual MDPs are powerful tools with wide applicability in areas from biostatistics to machine learning. However, specializing them to offline datasets has been challenging due to a lack of robust, theoretically backed methods. Our work tackles this problem by introducing a new approach towards adaptive estimation and cost optimization of contextual MDPs. This estimator, to the best of our knowledge, is the first of its kind, and is endowed with strong optimality guarantees. We achieve this by overcoming the key technical challenges evolving from the endogenous properties of contextual MDPs; such as non-stationarity, or model irregularity. Our guarantees are established under complete generality by utilizing the relatively recent and powerful statistical technique of $T$-estimation (Baraud, 2011). We first provide a procedure for selecting an estimator given a sample from a contextual MDP and use it to derive oracle risk bounds under two distinct, but nevertheless meaningful, loss functions. We then consider the problem of determining the optimal control with the aid of the aforementioned density estimate and provide finite sample guarantees for the cost function.

强化学习离线学习最优控制上下文MDP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。