解决强化学习中奖励缺失问题,提升决策准确性
Multi-armed Bandits with Missing Outcome
- 设计新算法应对随机与非随机缺失数据
- 在真实缺失场景下实现接近最优的后悔值
- 适用于医疗、推荐等缺失数据频发领域
尽管在线决策中的最小化后悔值算法已取得显著进展,但现实场景常引入额外复杂性,其中最严峻的是奖励缺失。忽视此问题或假设随机缺失会引发奖励估计偏差,导致线性后悔。尽管该问题具有重要实际意义,当前尚无系统方法处理非随机缺失。本文针对多臂赌博机中的缺失结果问题,分析不同缺失机制对可达到后悔界的影响,提出在缺失于随机(MAR)和非随机(MNAR)情形下均适用的算法。通过理论分析与仿真研究,证明考虑缺失机制能显著提升决策性能。
原文摘要 · Abstract (English)
While significant progress has been made in designing algorithms that minimize regret in online decision-making, real-world scenarios often introduce additional complexities, perhaps the most challenging of which is missing outcomes. Overlooking this aspect or simply assuming random missingness invariably leads to biased estimates of the rewards and may result in linear regret. Despite the practical relevance of this challenge, no rigorous methodology currently exists for systematically handling missingness, especially when the missingness mechanism is not random. In this paper, we address this gap in the context of multi-armed bandits (MAB) with missing outcomes by analyzing the impact of different missingness mechanisms on achievable regret bounds. We introduce algorithms that account for missingness under both missing at random (MAR) and missing not at random (MNAR) models. Through both analytical and simulation studies, we demonstrate the drastic improvements in decision-making by accounting for missingness in these settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。