arXiv:2603.19440stat.MLcs.LG2026-03

提出近似最优治疗策略集,让临床决策有更多可选方案。

Near-Equivalent Q-learning Policies for Dynamic Treatment Regimes

  • 引入ε容忍度,找出性能接近最优的多组治疗策略
  • 在模拟肿瘤治疗中验证了策略集的可行性与差异区域
  • 适合需要灵活备选方案的精准医疗场景

精准医学旨在根据个体特征定制治疗决策。这一目标通常通过动态治疗策略(Dynamic Treatment Regimes)建模,利用统计与机器学习方法生成随临床信息演变的序贯决策规则。现有方法大多仅输出每个阶段的单一最优治疗,导致唯一决策路径。然而在许多临床情境中,多种治疗方案可带来相似预期效果,仅关注单一最优策略可能掩盖有意义的替代选择。本文扩展了基于回顾性数据的Q-learning框架,引入由超参数ε控制的最差值容忍度,规定可接受的最优期望值最大偏差。该方法不再寻找单一最优策略,而是构建一组ε-最优策略,其表现保持在最优值的可控邻域内。这一形式将Q-learning从向量表示转变为矩阵表示,允许在逆向递推过程中共存多个可接受的价值函数。该方法生成一系列近似等效的治疗策略,并明确识别出多个决策表现相近的治疗无差别区域。我们在两个场景中验证该框架:一个单阶段问题揭示决策边界附近的无差别区域;一个基于模拟肿瘤模型的多阶段决策过程,描述肿瘤大小与治疗毒性动态变化。

原文摘要 · Abstract (English)

Precision medicine aims to tailor therapeutic decisions to individual patient characteristics. This objective is commonly formalized through dynamic treatment regimes, which use statistical and machine learning methods to derive sequential decision rules adapted to evolving clinical information. In most existing formulations, these approaches produce a single optimal treatment at each stage, leading to a unique decision sequence. However, in many clinical settings, several treatment options may yield similar expected outcomes, and focusing on a single optimal policy may conceal meaningful alternatives. We extend the Q-learning framework for retrospective data by introducing a worst-value tolerance criterion controlled by a hyperparameter $\varepsilon$, which specifies the maximum acceptable deviation from the optimal expected value. Rather than identifying a single optimal policy, the proposed approach constructs sets of $\varepsilon$-optimal policies whose performance remains within a controlled neighborhood of the optimum. This formulation shifts Q-learning from a vector-valued representation to a matrix-valued one, allowing multiple admissible value functions to coexist during backward recursion. The approach yields families of near-equivalent treatment strategies and explicitly identifies regions of treatment indifference where several decisions achieve comparable outcomes. We illustrate the framework in two settings: a single-stage problem highlighting indifference regions around the decision boundary, and a multi-stage decision process based on a simulated oncology model describing tumor size and treatment toxicity dynamics.

动态治疗精准医疗强化学习ε-最优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。