用短期数据预测长期政策效果,解决新疗法评估难题
Predicting Long Term Sequential Policy Value Using Softer Surrogates
- 基于短期替代指标与历史数据融合建模
- 仅需10%完整周期数据即可准确预测长期效果
- 适合医疗等长周期决策场景的政策评估
离策略策略评估(OPE)利用历史数据估算新策略的成效。然而,现有方法无法处理新策略引入全新行为的情况,这在医疗等领域尤为常见,如新药研发。由于长期结果需长时间观测(如多年临床试验),获取新策略的在线数据成本高昂。本文提出一种方法,通过短期替代指标与长期历史数据结合,在满足一定代理条件时,可准确预测新策略的长期价值。在艾滋病和败血症管理两个模拟医疗案例中,仅需观察完整时间跨度10%的数据,我们的估计器即可实现高精度预测。同时提供了双重稳健估计器的有限样本分析。
原文摘要 · Abstract (English)
Off-policy policy evaluation (OPE) estimates the outcome of a new policy using historical data collected from a different policy. However, existing OPE methods cannot handle cases when the new policy introduces novel actions. This issue commonly occurs in real-world domains, like healthcare, as new drugs and treatments are continuously developed. Novel actions necessitate on-policy data collection, which can be burdensome and expensive if the outcome of interest takes a substantial amount of time to observe--for example, in multi-year clinical trials. This raises a key question of how to predict the long-term outcome of a policy after only observing its short-term effects? Though in general this problem is intractable, under some surrogacy conditions, the short-term on-policy data can be combined with the long-term historical data to make accurate predictions about the new policy's long-term value. In two simulated healthcare examples--HIV and sepsis management--we show that our estimators can provide accurate predictions about the policy value only after observing 10\% of the full horizon data. We also provide finite sample analysis of our doubly robust estimators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。