arXiv:2507.11780econ.EMcs.LG2025-07被引 1

用软最大平滑法实现最优治疗策略价值的精准推断。

Inference on Optimal Policy Values and Other Irregular Functionals via Softmax Smoothing

  • 通过软最大平滑构建新估计器,避免非可导性难题。
  • 仅需固定次数拟合辅助模型,样本量增大时仍高效。
  • 适用于静态/动态治疗方案,支持复杂参数估计。

在因果推断中,构造未知最优治疗策略价值的置信区间是一项基础问题。理解最优策略价值有助于设计最大化收益的个性化治疗方案。然而,由于定义最优值的函数不可导,传统半参数推断方法无法直接应用。现有许多工作通过假设治疗无响应概率为零来规避这一问题,这在现实中不成立;而未做此假设的方法则需按样本量反复重拟合辅助模型。本文提出一种基于软最大平滑的简单估计器,适用于静态与动态治疗策略,仅需固定次数拟合辅助模型,且在无治疗非响应情况下具有统计效率。该方法无需半参数假设,但可在存在时加以利用。进一步证明,该平滑方法可用于估计由包含辅助成分的最大得分定义的一般参数,如条件Balke-Pearl界和L¹校准误差等。

原文摘要 · Abstract (English)

Constructing confidence intervals for the value of an (unknown) optimal treatment policy is a fundamental problem in causal inference. Insight into the optimal policy value can guide the development of reward-maximizing, individualized treatment regimes. However, because the functional that defines the optimal value is non-differentiable, standard semi-parametric approaches for performing inference fail to be directly applicable. Many existing works circumvent non-differentiability by making the unrealistic assumption of zero probability of treatment non-response, i.e. that every unit responds (either positively or negatively) to an assigned treatment. Further, works that don't circumvent this restriction rely on refitting nuisance models a number of times proportional to the sample size. In this paper, we construct and analyze a simple, softmax smoothing-based estimator for the value of an optimal treatment policy. Our estimator applies in both static and dynamic treatment regimes, only requires fitting a constant number of nuisance models, and is statistically efficient when there is zero probability of non-response to treatment. Also, while our estimator does not require making semi-parametric restrictions, it can exploit them when they exist. We further show how our softmax smoothing approach can be used to estimate general parameters that are specified as a maximum of scores involving nuisance components, and look at conditional Balke and Pearl bounds and $L^1$ calibration error as salient examples.

因果推断治疗策略软最大平滑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。