用用户主动提问加速多目标推荐的偏好学习,提升决策效率。
Provably Efficient Personalized Multi-Objective Bandits with Proactive Conversational Queries

- 通过用户提问构建结构化偏好信号,结合上下文建模
- 理论证明提问可加快偏好估计,降低后悔率
- 对噪声提问有鲁棒性,适合真实交互场景
在多目标强化学习中,个性化决策需学习用户在不同目标间的权衡。现有方法仅从结果反馈推断偏好,将偏好学习与奖励探索耦合。然而,用户常通过主动提问(如“便宜又干净的酒店”)揭示优先级,这类结构化信号未被利用。本文提出基于主动提问的框架,使用Plackett-Luce子集选择模型建模提问行为,发现仅靠提问无法准确学习偏好,存在固有的平移不变性障碍。为此,提出MO-PQUCB算法,融合提问锚定与带反馈的双探索机制,通过平移不变正则化实现混合学习。理论上证明,主动提问能加速偏好估计并改进后悔率。当提问受干扰时,进一步刻画统计极限并设计近似最优的鲁棒估计器。实验验证了理论与实际优势。
原文摘要 · Abstract (English)
Personalized decision-making in multi-objective bandits requires learning user-specific trade-offs among competing objectives. Since arm utility depends on both unknown rewards and unknown preferences, existing methods infer preferences only from utility feedback, entangling preference learning with reward exploration. In practice, however, users often reveal their priorities through proactive conversational queries (e.g., "cheap and clean hotel"), yet this structured signal is not leveraged. We formalize a proactive query-based framework in which user queries provide structured preference signals. Modeling these signals via a Plackett-Luce subset choice model, we show that query-only learning is insufficient due to a fundamental shift-invariance barrier. To resolve this, we introduce MO-PQUCB, a hybrid algorithm that integrates query-based preference anchoring with bandit feedback through shift-invariant regularization and dual-exploration UCB. We prove that proactive queries accelerate preference estimation and yield improved regret scaling over prior preference-aware MO-MAB methods. Under corrupted queries, we further characterize statistical limits and design a robust estimator achieving near-optimal performance when the corruption is sparse. Experiments validate both theoretical and practical gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。