arXiv:2605.15620stat.MLcs.LG2026-05

提出统一框架,让离线决策更安全,风险可控且无需额外数据成本。

Pessimistic Risk-Aware Policy Learning in Contextual Bandits

  • 构建分布式框架,统一优化多种风险度量
  • 理论证明统计误差率达最优 $ ilde{ m O}(1/ oot{n})$
  • 适合高风险场景的离线策略学习,如医疗、金融

我们研究风险感知的离线策略学习,旨在从记录数据中学习一种在广义风险准则下最优的决策规则。该问题在无法在线交互的高风险领域至关重要,需严格控制不利结果。然而,现有离线上下文老虎机研究或仅关注期望回报,或仅将风险考虑限于评估而非优化。本文提出一个统一的分布式框架,用于优化Lipschitz连续风险泛函,涵盖均值-方差、熵风险、条件风险价值等广泛风险度量。通过为基于重要性采样的分布估计器建立新颖的样本浓度不等式,我们的分析在不依赖严苛重叠假设的前提下,推导出数据相关的次优性界,达到 $ ilde{ m O}(1/ oot{n})$ 的收敛速率。该速率是极小极大最优的,与风险中性离线策略优化一致,表明优化一般Lipschitz风险准则不会带来额外的统计代价。

原文摘要 · Abstract (English)

We study risk-aware offline policy learning, aiming to learn a decision rule from logged data that is optimal under general risk criteria. This problem is crucial in high-stakes domains where online interaction is infeasible and adverse outcomes must be carefully controlled. However, existing literature on offline contextual bandits either centers on expected-reward criteria or restricts risk considerations to policy evaluation instead of optimization. In this work, we propose a unified distributional framework for optimizing Lipschitz-continuous risk functionals, a broad class of risk measures encompassing mean-variance, entropic risk, and conditional value-at-risk, among others. By developing novel empirical concentration inequalities for importance sampling-based distributional estimators, our analysis derives data-dependent suboptimality bounds with an $\tilde{\mathcal{O}}(1/\sqrt{n})$ rate, without relying on restrictive uniform overlap assumptions. This rate is minimax optimal and matches that of risk-neutral offline policy optimization, indicating that optimizing general Lipschitz risk criteria incurs no additional statistical cost relative to the expected-reward.

风险感知离线学习带约束优化强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。