arXiv:2412.20638cs.AIcs.LG2024-12

用短期数据预测长期政策效果,解决新疗法评估难题

Predicting Long Term Sequential Policy Value Using Softer Surrogates

  • 基于短期替代指标与历史数据融合建模
  • 仅需10%完整周期数据即可准确预测长期效果
  • 适合医疗等长周期决策场景的政策评估

离策略策略评估(OPE)利用历史数据估算新策略的成效。然而,现有方法无法处理新策略引入全新行为的情况,这在医疗等领域尤为常见,如新药研发。由于长期结果需长时间观测(如多年临床试验),获取新策略的在线数据成本高昂。本文提出一种方法,通过短期替代指标与长期历史数据结合,在满足一定代理条件时,可准确预测新策略的长期价值。在艾滋病和败血症管理两个模拟医疗案例中,仅需观察完整时间跨度10%的数据,我们的估计器即可实现高精度预测。同时提供了双重稳健估计器的有限样本分析。

原文摘要 · Abstract (English)

Off-policy policy evaluation (OPE) estimates the outcome of a new policy using historical data collected from a different policy. However, existing OPE methods cannot handle cases when the new policy introduces novel actions. This issue commonly occurs in real-world domains, like healthcare, as new drugs and treatments are continuously developed. Novel actions necessitate on-policy data collection, which can be burdensome and expensive if the outcome of interest takes a substantial amount of time to observe--for example, in multi-year clinical trials. This raises a key question of how to predict the long-term outcome of a policy after only observing its short-term effects? Though in general this problem is intractable, under some surrogacy conditions, the short-term on-policy data can be combined with the long-term historical data to make accurate predictions about the new policy's long-term value. In two simulated healthcare examples--HIV and sepsis management--we show that our estimators can provide accurate predictions about the policy value only after observing 10\% of the full horizon data. We also provide finite sample analysis of our doubly robust estimators.

策略评估医疗决策替代指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。