arXiv:2503.12760stat.MLcs.LG2025-03

提出新方法同时学习与评估安全策略,提升多目标干预的可靠性。

SNPL: Simultaneous Policy Learning and Evaluation for Safe Multi-Objective Policy Improvement

  • 利用算法稳定性实现无需数据分割的联合策略学习与评估
  • 在小样本低信噪比下检测率提升300%,策略收益提升150%
  • 适合需保障安全约束的个性化数字干预场景

为设计有效的数字干预措施,研究者面临如何基于离线数据学习平衡多目标的决策策略的挑战。通常目标是最大化目标结果,同时确保保护性结果不出现不良变化。为提供可信建议,研究者不仅需识别满足目标与保护性结果变化期望的策略,还需对这些策略引发的变化提供概率保证。然而,实际中策略空间通常庞大,数字实验数据常存在相对于噪声的微弱效应。此时,标准方法如数据分割或多假设检验常导致策略选择不稳定或统计功效不足。本文提出安全噪声策略学习(SNPL),借助算法稳定性概念解决上述挑战。该方法可在不进行数据分割的情况下,利用全部数据同时完成策略学习与高置信度保证。我们提出了有限样本与渐近版本的算法,确保推荐策略在避免保护性结果恶化或实现目标结果提升方面具有高概率保证。我们在真实世界中个性化短信推送的应用上测试了两种变体。实证结果表明,在大策略空间和低信噪比条件下,本方法在有限样本与渐近安全保证上均有显著提升,检测率最高提升300%,策略收益提升150%,且所需样本量显著更小。

原文摘要 · Abstract (English)

To design effective digital interventions, experimenters face the challenge of learning decision policies that balance multiple objectives using offline data. Often, they aim to develop policies that maximize goal outcomes, while ensuring there are no undesirable changes in guardrail outcomes. To provide credible recommendations, experimenters must not only identify policies that satisfy the desired changes in goal and guardrail outcomes, but also offer probabilistic guarantees about the changes these policies induce. In practice, however, policy classes are often large, and digital experiments tend to produce datasets with small effect sizes relative to noise. In this setting, standard approaches such as data splitting or multiple testing often result in unstable policy selection and/or insufficient statistical power. In this paper, we provide safe noisy policy learning (SNPL), a novel approach that leverages the concept of algorithmic stability to address these challenges. Our method enables policy learning while simultaneously providing high-confidence guarantees using the entire dataset, avoiding the need for data-splitting. We present finite-sample and asymptotic versions of our algorithm that ensure the recommended policy satisfies high-probability guarantees for avoiding guardrail regressions and/or achieving goal outcome improvements. We test both variants of our approach approach empirically on a real-world application of personalizing SMS delivery. Our results on real-world data suggest that our approach offers dramatic improvements in settings with large policy classes and low signal-to-noise across both finite-sample and asymptotic safety guarantees, offering up to 300\% improvements in detection rates and 150\% improvements in policy gains at significantly smaller sample sizes.

策略学习安全优化多目标离线强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。