解决离策略评估中的不确定性与因果难题,提升决策可靠性
Uncertainty Quantification and Causal Considerations for Off-Policy Decision Making
- 提出边际比率估计算法,降低重要性采样方差
- 构建置信区间框架,实现有限样本下的不确定性量化
- 引入可证伪性检验,识别数字孪生模型失效场景
离策略评估(OPE)是稳健决策的核心挑战,旨在利用不同策略收集的数据评估新策略表现。现有方法在统计不确定性和因果考量上存在局限。本文提出三项工作:首先,针对重要性采样估计器方差过高问题,提出边际比率(MR)估计算法,通过关注结果的边缘分布而非直接政策转移,提升上下文老虎机中的鲁棒性;其次,提出合约定价离策略预测(COPP),一种基于原则的不确定性量化方法,提供有限样本下的预测区间,适用于风险敏感场景;最后,针对序列决策中因果不可识别问题,建立在任意未测量混杂条件下仍有效的新型边界,并应用于数字孪生模型可靠性评估,引入可证伪性框架以识别模型预测与真实行为偏离的场景。研究为不确定环境下的稳健决策提供了新视角和严谨方法。
原文摘要 · Abstract (English)
Off-policy evaluation (OPE) is a critical challenge in robust decision-making that seeks to assess the performance of a new policy using data collected under a different policy. However, the existing OPE methodologies suffer from several limitations arising from statistical uncertainty as well as causal considerations. In this thesis, we address these limitations by presenting three different works. Firstly, we consider the problem of high variance in the importance-sampling-based OPE estimators. We introduce the Marginal Ratio (MR) estimator, a novel OPE method that reduces variance by focusing on the marginal distribution of outcomes rather than direct policy shifts, improving robustness in contextual bandits. Next, we propose Conformal Off-Policy Prediction (COPP), a principled approach for uncertainty quantification in OPE that provides finite-sample predictive intervals, ensuring robust decision-making in risk-sensitive applications. Finally, we address causal unidentifiability in off-policy decision-making by developing novel bounds for sequential decision settings, which remain valid under arbitrary unmeasured confounding. We apply these bounds to assess the reliability of digital twin models, introducing a falsification framework to identify scenarios where model predictions diverge from real-world behaviour. Our contributions provide new insights into robust decision-making under uncertainty and establish principled methods for evaluating policies in both static and dynamic settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。