arXiv:2510.03912cs.LG2025-10

针对医疗数据中的簇内相关性,提出改进的强化学习算法,显著降低决策误差。

Generalized Fitted Q-Iteration with Clustered Data

  • 将广义估计方程融入拟合Q迭代,处理数据簇内相关性
  • 实验显示平均减少50%的后悔值,优于标准FQI
  • 适合医疗、移动健康等存在分组数据的应用场景

本文研究具有簇结构数据的强化学习问题,常见于医疗应用场景。提出一种改进的广义拟合Q迭代(Generalized FQI)算法,通过引入广义估计方程来建模簇内相关性。理论上证明:当相关结构正确指定时,所提的Q函数和策略估计器达到最优;当结构误设时,二者仍具一致性。通过模拟实验与移动健康数据集分析发现,该方法相比标准FQI平均可减少50%的后悔值。

原文摘要 · Abstract (English)

This paper focuses on reinforcement learning (RL) with clustered data, which is commonly encountered in healthcare applications. We propose a generalized fitted Q-iteration (FQI) algorithm that incorporates generalized estimating equations into policy learning to handle the intra-cluster correlations. Theoretically, we demonstrate (i) the optimalities of our Q-function and policy estimators when the correlation structure is correctly specified, and (ii) their consistencies when the structure is mis-specified. Empirically, through simulations and analyses of a mobile health dataset, we find the proposed generalized FQI achieves, on average, a half reduction in regret compared to the standard FQI.

强化学习医疗AI簇数据决策优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。