为安全关键的决策问题提供基于风险等级的量化下界分析方法
Risk level dependent Minimax Quantile lower bounds for Interactive Statistical Decision Making
- 构建交互式统计决策框架下的高概率Fano与Le Cam工具
- 首次获得风险等级明确的极小极大分位数下界,含分位数到期望的转换
- 适用于带风险控制需求的强化学习与贝叶斯优化场景
极小极大风险和遗憾关注期望值,忽略了安全关键的带宽和强化学习中罕见失败的影响。极小极大分位数能捕捉这些尾部事件。现有研究有三个方向:非交互估计中的极小极大分位数约束;统一的交互分析聚焦期望风险而非特定风险等级的分位数边界;高概率带宽边界仍缺乏针对一般交互协议的分位数专用工具。为填补这一空白,在交互式统计决策框架中,我们开发了高概率Fano和Le Cam工具,推导出风险等级明确的极小极大分位数下界,包括分位数到期望的转换以及严格与下极小极大分位数之间的紧密联系。对两臂高斯带宽问题的实例化立即恢复了最优率边界。
原文摘要 · Abstract (English)
Minimax risk and regret focus on expectation, missing rare failures critical in safety-critical bandits and reinforcement learning. Minimax quantiles capture these tails. Three strands of prior work motivate this study: minimax-quantile bounds restricted to non-interactive estimation; unified interactive analyses that focus on expected risk rather than risk level specific quantile bounds; and high-probability bandit bounds that still lack a quantile-specific toolkit for general interactive protocols. To close this gap, within the interactive statistical decision making framework, we develop high-probability Fano and Le Cam tools and derive risk level explicit minimax-quantile bounds, including a quantile-to-expectation conversion and a tight link between strict and lower minimax quantiles. Instantiating these results for the two-armed Gaussian bandit immediately recovers optimal-rate bounds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。