arXiv:2604.22140stat.MLcs.LG2026-04

用影响函数梯度优化带分布效用的强化学习,提升决策鲁棒性。

Concave Statistical Utility Maximization Bandits via Influence-Function Gradients

  • 基于影响函数推导梯度,通过乘法权重更新实现优化。
  • 理论证明了可分离的优化误差与估计偏差,保证收敛性。
  • 适用于方差、Wasserstein等分布目标,适合风险敏感场景。

我们研究一类新型多臂老虎机问题,其目标是长期奖励分布的统计泛函(如方差、Wasserstein距离),而非仅期望收益。在温和连续性假设下,无限时域问题可转化为对平稳混合策略的优化:每个定义在单纯形上的权重向量 $w$ 对应一个混合分布 $P^w$,性能由凹效用函数 $U(w)=\mathfrak U(P^w)$ 衡量。对于可微的统计效用,我们利用影响函数微积分,从老虎机反馈中构建随机梯度估计器。由此提出一种在截断单纯形上的熵镜面上升算法,通过乘法权重更新和影响函数的插值估计实现。我们建立了可分离的后悔界,将镜面上升的优化误差与影响函数估计带来的偏差区分开。该框架适用于一般凹分布效用,并以方差和Wasserstein目标为例进行了数值实验,对比了精确与插值实现的效果。

原文摘要 · Abstract (English)

We study stochastic multi-armed bandits in which the objective is a statistical functional of the long-run reward distribution, rather than expected reward alone. Under mild continuity assumptions, we show that the infinite-horizon problem reduces to optimizing over stationary mixed policies: each weight vector \(w\) on the simplex induces a mixture law \(P^w\), and performance is measured by the concave utility \(U(w)=\mathfrak U(P^w)\). For differentiable statistical utilities, we use influence-function calculus to derive stochastic gradient estimators from bandit feedback. This leads to an entropic mirror-ascent algorithm on a truncated simplex, implemented through multiplicative-weights updates and plug-in estimates of the influence function. We establish regret bounds that separate the mirror-ascent optimization error from the bias caused by estimating the influence function. The framework is developed for general concave distributional utilities and illustrated through variance and Wasserstein objectives, with numerical experiments comparing exact and plug-in influence-function implementations.

强化学习带分布目标梯度优化风险敏感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。