在分布偏移下,用偏好锥优化上下文老虎机的决策效率
Vector preference-based contextual bandits under distributional shifts
- 基于偏好锥设计自适应离散化与乐观淘汰策略
- 提出偏好型后悔度衡量指标,可量化帕累托前沿差距
- 在分布偏移下仍保持良好性能,适合多目标决策场景
研究在奖励向量按给定偏好锥排序的条件下,存在分布偏移时的上下文老虎机学习问题。提出一种基于自适应离散化与乐观淘汰的策略,能自动适应潜在的分布偏移。为评估该策略性能,引入偏好型后悔度,以帕累托前沿间的距离衡量政策表现。在不同分布偏移假设下建立了后悔上界,其结果推广了无分布偏移和向量奖励情形的已有结论,并在分布偏移存在时仍能随问题参数平滑缩放。
原文摘要 · Abstract (English)
We consider contextual bandit learning under distribution shift when reward vectors are ordered according to a given preference cone. We propose an adaptive-discretization and optimistic elimination based policy that self-tunes to the underlying distribution shift. To measure the performance of this policy, we introduce the notion of preference-based regret which measures the performance of a policy in terms of distance between Pareto fronts. We study the performance of this policy by establishing upper bounds on its regret under various assumptions on the nature of distribution shift. Our regret bounds generalize known results for the existing case of no distribution shift and vectorial reward settings, and scale gracefully with problem parameters in presence of distribution shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。