针对用户偏好定制的多目标强化学习算法,更贴合真实场景
Provably Efficient Multi-Objective Bandit Algorithms under Preference-Centric Customization
- 根据用户偏好在帕累托前沿内优化选择
- 提出两种新算法,理论证明近似最优后悔率
- 适合需要个性化推荐与决策的系统设计
多目标多臂赌博机(MO-MAB)传统上追求帕累托最优。但在现实场景中,用户对各目标的偏好各异,导致一个帕累托最优解可能对某些用户表现极差。为此,本文研究在显式用户偏好下的偏好感知多目标多臂赌博机框架,将优化焦点从追求帕累托最优转向在帕累托前沿内进行偏好中心化定制优化。这是首个对显式用户偏好下定制化多目标优化进行理论分析的研究。基于实际应用需求,本文探讨了偏好未知和偏好隐藏两种情形,分别面临不同算法设计与分析挑战。核心机制包括偏好估计与偏好感知优化,通过创新分析技术,建立了所提算法的近似最优后悔界。实验验证了方法的有效性。
原文摘要 · Abstract (English)
Multi-objective multi-armed bandit (MO-MAB) problems traditionally aim to achieve Pareto optimality. However, real-world scenarios often involve users with varying preferences across objectives, resulting in a Pareto-optimal arm that may score high for one user but perform quite poorly for another. This highlights the need for customized learning, a factor often overlooked in prior research. To address this, we study a preference-aware MO-MAB framework in the presence of explicit user preference. It shifts the focus from achieving Pareto optimality to further optimizing within the Pareto front under preference-centric customization. To our knowledge, this is the first theoretical study of customized MO-MAB optimization with explicit user preferences. Motivated by practical applications, we explore two scenarios: unknown preference and hidden preference, each presenting unique challenges for algorithm design and analysis. At the core of our algorithms are preference estimation and preference-aware optimization mechanisms to adapt to user preferences effectively. We further develop novel analytical techniques to establish near-optimal regret of the proposed algorithms. Strong empirical performance confirm the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。