arXiv:2602.01825stat.MEcs.LG2026-02

从多机构数据中学习鲁棒序列决策,避免对强假设的依赖。

Learning Sequential Decisions from Multiple Sources via Group-Robust Markov Decision Processes

  • 基于共享特征的线性结构建模跨机构差异,保持可计算性。
  • 通过逐特征最坏情况聚合与数据相关惩罚项,提升策略稳健性。
  • 适用于医疗等异构数据场景,尤其适合有先验相似性知识的领域。

我们常从多个站点(如医院)收集数据,这些数据具有共同结构但存在异质性。本文旨在从此类离线、多站点数据集中学习鲁棒的序列决策策略。为建模跨站点不确定性,我们研究具有组-线性结构的分布鲁棒马尔可夫决策过程:所有站点共享一个公共特征映射,且转移核与期望奖励函数均在此共享特征上线性。引入特征级(d-矩形)不确定性集,在保持可计算的鲁棒贝尔曼递推的同时,保留关键跨站点结构。基于此,我们提出一种基于悲观值迭代的离线算法,包含:(i) 每站点的岭回归用于贝尔曼目标估计,(ii) 特征级最坏情况(行级最小化)聚合,(iii) 基于设计矩阵逆对角线计算的数据依赖悲观惩罚项。我们进一步提出聚类级扩展,利用站点相似性先验合并相似站点以提高样本效率。在鲁棒部分覆盖假设下,证明了所得策略的次优性界。整体框架解决了异构数据源下的多站点学习问题,并提供了一种无需强状态-动作矩形性假设的鲁棒规划方法。

原文摘要 · Abstract (English)

We often collect data from multiple sites (e.g., hospitals) that share common structure but also exhibit heterogeneity. This paper aims to learn robust sequential decision-making policies from such offline, multi-site datasets. To model cross-site uncertainty, we study distributionally robust MDPs with a group-linear structure: all sites share a common feature map, and both the transition kernels and expected reward functions are linear in these shared features. We introduce feature-wise (d-rectangular) uncertainty sets, which preserve tractable robust Bellman recursions while maintaining key cross-site structure. Building on this, we then develop an offline algorithm based on pessimistic value iteration that includes: (i) per-site ridge regression for Bellman targets, (ii) feature-wise worst-case (row-wise minimization) aggregation, and (iii) a data-dependent pessimism penalty computed from the diagonals of the inverse design matrices. We further propose a cluster-level extension that pools similar sites to improve sample efficiency, guided by prior knowledge of site similarity. Under a robust partial coverage assumption, we prove a suboptimality bound for the resulting policy. Overall, our framework addresses multi-site learning with heterogeneous data sources and provides a principled approach to robust planning without relying on strong state-action rectangularity assumptions.

强化学习多源学习鲁棒决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。