arXiv:2509.09772cs.LGstat.AP2025-09

用分层风险控制提升医保人群健康管理的公平性与安全性。

Hybrid Adaptive Conformal Offline Reinforcement Learning for Fair Population Health Management

  • 分离风险校准与偏好优化,动态屏蔽高危动作
  • 在277万条决策中实现0.81的危险事件预测AUC
  • 支持按年龄/性别/种族审计,适合医疗决策系统

针对医保人群的健康管理需兼顾安全、公平与可审计性。本文提出混合自适应共形离线强化学习(HACO)框架,将风险校准与偏好优化分离,在每步决策中选择协调动作(如联系对象、沟通方式、是否转介专业服务),同时控制近端不良利用事件(如急诊或住院)风险。基于包含277万条序列决策的Waymark去标识化数据集(覆盖168,126名患者),HACO(i)训练轻量级不良事件风险模型,(ii)通过共形推断确定目标风险水平下的阈值(τ ~0.038,α = 0.10),屏蔽不安全动作,(iii)在安全子集上学习偏好策略。采用版本无关拟合Q评估(FQE)在分层子集上评估策略,并按年龄、性别、种族审计表现。结果表明HACO实现强风险区分能力(AUC ~0.81),阈值校准良好且保持高安全覆盖率。子群分析显示不同人口统计学群体间价值估计存在系统性差异,凸显公平审计重要性。实证表明共形风险门控可无缝融入离线强化学习,为健康管理团队提供保守、可审计的决策支持。

原文摘要 · Abstract (English)

Population health management programs for Medicaid populations coordinate longitudinal outreach and services (e.g., benefits navigation, behavioral health, social needs support, and clinical scheduling) and must be safe, fair, and auditable. We present a Hybrid Adaptive Conformal Offline Reinforcement Learning (HACO) framework that separates risk calibration from preference optimization to generate conservative action recommendations at scale. In our setting, each step involves choosing among common coordination actions (e.g., which member to contact, by which modality, and whether to route to a specialized service) while controlling the near-term risk of adverse utilization events (e.g., unplanned emergency department visits or hospitalizations). Using a de-identified operational dataset from Waymark comprising 2.77 million sequential decisions across 168,126 patients, HACO (i) trains a lightweight risk model for adverse events, (ii) derives a conformal threshold to mask unsafe actions at a target risk level, and (iii) learns a preference policy on the resulting safe subset. We evaluate policies with a version-agnostic fitted Q evaluation (FQE) on stratified subsets and audit subgroup performance across age, sex, and race. HACO achieves strong risk discrimination (AUC ~0.81) with a calibrated threshold ( τ ~0.038 at α = 0.10), while maintaining high safe coverage. Subgroup analyses reveal systematic differences in estimated value across demographics, underscoring the importance of fairness auditing. Our results show that conformal risk gating integrates cleanly with offline RL to deliver conservative, auditable decision support for population health management teams.

强化学习医疗决策公平性风险控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。