多目标推荐中,多个优质选项能带来隐式探索优势。
Blessings of Multiple Good Arms in Multi-Objective Linear Bandits
- 利用多个优质选项实现无需显式探索的算法设计
- 贪心策略在多数回合下仍能达成优异性能
- 适用于追求公平性的多目标决策场景
多目标强化学习传统上被认为比单目标更复杂,需同时优化多个目标。然而,当多个目标均存在多个优质动作时,反而会引发一种意外优势——隐式探索。在此条件下,我们证明:即使大多数轮次采用贪心选择,简单算法依然能在理论和实验上表现出色。据我们所知,这是首个在无上下文分布假设的前提下,将隐式探索引入多目标与参数化强化学习设置的研究。此外,我们提出了一个有效的帕累托公平性分析框架,为多目标强化学习算法的公平性提供严谨评估方法。
原文摘要 · Abstract (English)
The multi objective bandit setting has traditionally been regarded as more complex than the single objective case, as multiple objectives must be optimized simultaneously. In contrast to this prevailing view, we demonstrate that when multiple good arms exist for multiple objectives, they can induce a surprising benefit, implicit exploration. Under this condition, we show that simple algorithms that greedily select actions in most rounds can nonetheless achieve strong performance, both theoretically and empirically. To our knowledge, this is the first study to introduce implicit exploration in both multi objective and parametric bandit settings without any distributional assumptions on the contexts. We further introduce a framework for effective Pareto fairness, which provides a principled approach to rigorously analyzing fairness of multi objective bandit algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。