arXiv:2510.04728cs.LG2025-10

在风险敏感的老虎机问题中,提出最优选最优臂的算法。

EVaR-Optimal Arm Identification in Bandits

  • 基于熵风险价值准则设计选最优臂算法。
  • 首次建立该问题下样本复杂度的理论下界并实现渐近匹配。
  • 适合金融等高风险场景的稳健决策研究者参考。

我们研究了多臂老虎机框架下基于熵风险价值(EVaR)准则的固定置信度最优臂识别(BAI)问题。分析采用非参数设定,允许奖励分布为[0,1]区间内的任意有界分布。该公式解决了金融等高风险环境中对风险规避决策的需求,超越了简单的期望值优化。我们提出一种δ-正确的、基于追踪-停止(Track-and-Stop)的算法,并推导出相应的期望样本复杂度下界,证明其渐近可达到。算法实现与下界刻画均需求解一个复杂的凸优化问题及一个相关的较简单非凸问题。

原文摘要 · Abstract (English)

We study the fixed-confidence best arm identification (BAI) problem within the multi-armed bandit (MAB) framework under the Entropic Value-at-Risk (EVaR) criterion. Our analysis considers a nonparametric setting, allowing for general reward distributions bounded in [0,1]. This formulation addresses the critical need for risk-averse decision-making in high-stakes environments, such as finance, moving beyond simple expected value optimization. We propose a $δ$-correct, Track-and-Stop based algorithm and derive a corresponding lower bound on the expected sample complexity, which we prove is asymptotically matched. The implementation of our algorithm and the characterization of the lower bound both require solving a complex convex optimization problem and a related, simpler non-convex one.

强化学习风险建模贝叶斯优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。