在有限预算和置信度下,找最优风险收益组合的强化学习方法
Risk-Averse Best Arm Set Identification with Fixed Budget and Fixed Confidence
- 统一框架适配固定预算与固定置信两种场景
- 能准确识别风险收益权衡下的最优动作集
- 适合需兼顾收益与稳定性的实际决策任务
在不确定环境中的决策问题,常需在最大化期望收益的同时最小化风险。本文提出一种新的随机老虎机优化问题设置,联合考虑期望回报与关联不确定性,以均值-方差(Mean-Variance, MV)准则量化风险。不同于仅关注期望回报的传统老虎机模型,本研究目标是高效且准确地识别帕累托最优的动作集,实现期望性能与风险之间的最佳权衡。我们提出一个统一的元算法框架,通过自适应设计置信区间,在固定预算与固定置信两种场景下均能使用相同的样本探索策略。理论证明了两种设置下返回解的正确性。此外,在合成基准上的大量实验表明,该方法在准确性和样本效率上均优于现有方法,凸显其在不确定环境中进行风险敏感决策的广泛适用性。
原文摘要 · Abstract (English)
Decision making under uncertain environments in the maximization of expected reward while minimizing its risk is one of the ubiquitous problems in many subjects. Here, we introduce a novel problem setting in stochastic bandit optimization that jointly addresses two critical aspects of decision-making: maximizing expected reward and minimizing associated uncertainty, quantified via the mean-variance(MV) criterion. Unlike traditional bandit formulations that focus solely on expected returns, our objective is to efficiently and accurately identify the Pareto-optimal set of arms that strikes the best trade-off between expected performance and risk. We propose a unified meta-algorithmic framework capable of operating under both fixed-confidence and fixed-budget regimes, achieved through adaptive design of confidence intervals tailored to each scenario using the same sample exploration strategy. We provide theoretical guarantees on the correctness of the returned solutions in both settings. To complement this theoretical analysis, we conduct extensive empirical evaluations across synthetic benchmarks, demonstrating that our approach outperforms existing methods in terms of both accuracy and sample efficiency, highlighting its broad applicability to risk-aware decision-making tasks in uncertain environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。