动态环境中自适应调整风险水平的在线学习方法
Risk-Averse Learning with Varying Risk Levels
- 用函数变化和风险变化度量环境动态性,设计自适应算法
- 在有限采样下实现与环境变化相关的动态后悔界
- 适合安全关键场景中风险敏感的实时决策
在安全关键型决策中,环境可能随时间演变,学习者需相应调整风险水平。本文研究动态环境下风险规避的在线优化问题,采用条件风险价值(CVaR)作为风险度量。为刻画环境与风险水平的动态性,引入函数变差指标和新型风险水平变差指标。考虑两种信息设置:一阶情形(可获取函数值与梯度)和零阶情形(仅可获取函数评估)。针对两种情况,分别设计了在有限采样预算下的风险规避学习算法,并分析其动态后悔上界,结果以函数变差、风险水平变差及总采样数表示。理论分析表明算法在非平稳且风险敏感环境中具备良好适应性。最后通过数值实验验证方法有效性。
原文摘要 · Abstract (English)
In safety-critical decision-making, the environment may evolve over time, and the learner adjusts its risk level accordingly. This work investigates risk-averse online optimization in dynamic environments with varying risk levels, employing Conditional Value-at-Risk (CVaR) as the risk measure. To capture the dynamics of the environment and risk levels, we employ the function variation metric and introduce a novel risk-level variation metric. Two information settings are considered: a first-order scenario, where the learner observes both function values and their gradients; and a zeroth-order scenario, where only function evaluations are available. For both cases, we develop risk-averse learning algorithms with a limited sampling budget and analyze their dynamic regret bounds in terms of function variation, risk-level variation, and the total number of samples. The regret analysis demonstrates the adaptability of the algorithms in non-stationary and risk-sensitive settings. Finally, numerical experiments are presented to demonstrate the efficacy of the methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。