用贝叶斯框架优化决策规则,提升政策学习的稳定性与可解释性。
General Bayesian Policy Learning
- 基于损失函数的贝叶斯更新,将决策问题转化为最大化预期福利。
- 引入平方损失代理,使惩罚项控制下可等价求解收益差的误差最小化。
- 提出GBPLNet神经网络模型,适用于治疗选择与投资组合优化等场景。
本文提出一种通用贝叶斯政策学习框架,针对决策者从给定动作集中选择以最大化预期福利的问题,如治疗选择和投资组合优化。在此类问题中,统计目标是决策规则,而非预测每个潜在结果。通过基于损失的贝叶斯更新,采用平方损失代理实现福利最大化。我们证明:在策略类上对经验福利进行带二次惩罚(由调优参数 $ζ>0$ 控制)的最大化,等价于最小化结果差异的缩放平方误差。由此得到的决策规则广义贝叶斯后验具有两种等价表征:高斯伪似然形式与决策论损失基础表征。作为具体实现,引入GBPLNet——一种使用tanh压缩输出的神经网络。最后,建立了关于代理风险与相应正则化福利的PAC-Bayes型保证。
原文摘要 · Abstract (English)
This study proposes a General Bayes framework for policy learning. We consider decision problems in which a decision-maker chooses an action from a given set to maximize expected welfare. Typical examples include treatment choice and portfolio optimization. In such problems, the statistical target is a decision rule, and predicting each potential outcome is not necessarily of primary interest. We formulate this policy-learning problem through loss-based Bayesian updating. Our main technical device is a squared-loss surrogate for welfare maximization. We show that maximizing empirical welfare over a policy class with a quadratic penalty controlled by a tuning parameter $ζ>0$ is equivalent to minimizing a scaled squared error in the outcome difference. The resulting General Bayes posterior over decision rules admits two equivalent characterizations: a Gaussian pseudo-likelihood representation and a decision-theoretic loss-based characterization. As one implementation, we introduce GBPLNet, a neural network implementation with a tanh-squashed output. Finally, we establish PAC-Bayes-type guarantees for the surrogate risk and the corresponding penalized welfare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。