用轻量级强化学习优化医疗协调中的服务方式选择。
Test-Time Learning and Inference-Time Deliberation for Efficiency-First Offline Reinforcement Learning in Care Coordination and Population Health Management
- 测试时通过局部邻域校准调整策略,提升适应性。
- 推理时用小Q集合结合不确定性与成本,实现高效决策。
- 支持可审计、可调参,适合医疗系统落地应用。
医疗协调与人群健康管理需服务于大量医疗补助及安全网人群,要求可审计、高效且可适应。尽管干预的临床风险较低,但文本、电话、视频和面对面访问的时间与机会成本差异显著。本文提出一种轻量级离线强化学习方法,通过(i)测试时学习的局部邻域校准,以及(ii)推理时通过小型Q-ensemble融合预测不确定性与时间/努力成本的决策机制。该方法提供可解释的调节参数,如邻域大小与不确定性/成本惩罚,并保持可审计的训练流程。在去标识化真实运营数据集上评估,TTL+ITD实现了稳定的值估计,具备可预测的效率权衡与子群体审计能力。
原文摘要 · Abstract (English)
Care coordination and population health management programs serve large Medicaid and safety-net populations and must be auditable, efficient, and adaptable. While clinical risk for outreach modalities is typically low, time and opportunity costs differ substantially across text, phone, video, and in-person visits. We propose a lightweight offline reinforcement learning (RL) approach that augments trained policies with (i) test-time learning via local neighborhood calibration, and (ii) inference-time deliberation via a small Q-ensemble that incorporates predictive uncertainty and time/effort cost. The method exposes transparent dials for neighborhood size and uncertainty/cost penalties and preserves an auditable training pipeline. Evaluated on a de-identified operational dataset, TTL+ITD achieves stable value estimates with predictable efficiency trade-offs and subgroup auditing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。