解决延迟反馈下的风险规避学习问题,提升决策稳健性。
Risk-averse learning with delayed feedback
- 采用单点与两点零阶优化方法应对延迟反馈风险
- 在无延迟时达到现有最优零阶方法的误差边界
- 两点法比单点法更优,适合对风险敏感的动态决策场景
在真实场景中,风险规避学习有助于降低潜在负面结果。然而,延迟反馈使得风险评估与管理变得困难。本文研究以条件风险价值(CVaR)为风险度量的风险规避学习,并考虑具有随机但有界延迟的反馈。我们提出了两种基于单点和两点零阶优化的风险规避学习算法,分析了其动态遗憾,依赖于累计延迟和总采样次数。在无延迟情况下,遗憾界与已有零阶随机梯度方法一致。此外,两点算法优于单点算法,实现了更小的遗憾界。通过动态定价问题的数值实验验证了算法性能。
原文摘要 · Abstract (English)
In real-world scenarios, risk-averse learning is valuable for mitigating potential adverse outcomes. However, the delayed feedback makes it challenging to assess and manage risk effectively. In this paper, we investigate risk-averse learning using Conditional Value at Risk (CVaR) as risk measure, while incorporating feedback with random but bounded delays. We develop two risk-averse learning algorithms that rely on one-point and two-point zeroth-order optimization approaches, respectively. The dynamic regrets of the algorithms are analyzed in terms of the cumulative delay and the number of total samplings. In the absence of delay, the regret bounds match the established bounds of zeroth-order stochastic gradient methods for risk-averse learning. Furthermore, the two-point risk-averse learning outperforms the one-point algorithm by achieving a smaller regret bound. We provide numerical experiments on a dynamic pricing problem to demonstrate the performance of the algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。