用贝叶斯+树搜索高效获取用户偏好,少问几轮也能准
Preference Construction: A Bayesian Interactive Preference Elicitation Framework Based on Monte Carlo Tree Search
- 基于变分贝叶斯推断用户偏好,量化不确定性
- 用蒙特卡洛树搜索优化提问策略,减少总不确定度
- 适合需要少交互的决策辅助系统,如个性化推荐
我们提出一种新型偏好学习框架,可在有限交互轮次内高效捕捉参与者偏好。首先,采用变分贝叶斯方法通过估计后验分布来推断偏好模型,并管理信息有限带来的不确定性。其次,设计自适应提问策略,以最大化累积不确定性降低为目标,将提问建模为有限马尔可夫决策过程,利用蒙特卡洛树搜索优先选择有潜力的问题路径,兼顾长期效果并避免短视。第三,将该框架应用于多准则决策辅助,以成对比较作为偏好信息,加性价值函数作为偏好模型,并引入重参数化技巧缓解高方差问题,提升鲁棒性与效率。在真实世界和合成数据集上的计算实验表明,该框架在捕捉偏好和实现有限交互下的不确定性降低方面均优于基线方法。
原文摘要 · Abstract (English)
We present a novel preference learning framework to capture participant preferences efficiently within limited interaction rounds. It involves three main contributions. First, we develop a variational Bayesian approach to infer the participant's preference model by estimating posterior distributions and managing uncertainty from limited information. Second, we propose an adaptive questioning policy that maximizes cumulative uncertainty reduction, formulating questioning as a finite Markov decision process and using Monte Carlo Tree Search to prioritize promising question trajectories. By considering long-term effects and leveraging the efficiency of the Bayesian approach, the policy avoids shortsightedness. Third, we apply the framework to Multiple Criteria Decision Aiding, with pairwise comparison as the preference information and an additive value function as the preference model. We integrate the reparameterization trick to address high-variance issues, enhancing robustness and efficiency. Computational studies on real-world and synthetic datasets demonstrate the framework's practical usability, outperforming baselines in capturing preferences and achieving superior uncertainty reduction within limited interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。