用增强反馈提升稀疏场景下的偏好学习效率
Preference is More Than Comparisons: Rethinking Dueling Bandits with Augmented Human Feedback
- 通过增强反馈构建更可靠的置信区间,改进无模型双人赌博机框架
- 在推荐、多目标优化等任务中表现优于传统方法
- 适合需要高效人机交互的个性化系统开发
交互式偏好获取(IPE)旨在大幅降低人类标注成本,同时在个性化系统中收集人类偏好。双人赌博机(DB)算法基于成对比较实现最优决策,但在人类反馈稀疏时效率低下。现有方法依赖参数化奖励模型,其固定假设易受模型误设影响。本文提出基于反馈增强的新视角,改进无模型DB框架:引入广义集中性质下的增强置信边界,结合后悔分析揭示多因素性能权衡。原型算法在推荐、多目标优化及大模型响应优化等多个IPE基准上表现优异,验证了该方法在更广泛场景下实现可证明高效的潜力。
原文摘要 · Abstract (English)
Interactive preference elicitation (IPE) aims to substantially reduce human effort while acquiring human preferences in wide personalization systems. Dueling bandit (DB) algorithms enable optimal decision-making in IPE building on pairwise comparisons. However, they remain inefficient when human feedback is sparse. Existing methods address sparsity by heavily relying on parametric reward models, whose rigid assumptions are vulnerable to misspecification. In contrast, we explore an alternative perspective based on feedback augmentation, and introduce critical improvements to the model-free DB framework. Specifically, we introduce augmented confidence bounds to integrate augmented human feedback under generalized concentration properties, and analyze the multi-factored performance trade-off via regret analysis. Our prototype algorithm achieves competitive performance across several IPE benchmarks, including recommendation, multi-objective optimization, and response optimization for large language models, demonstrating the potential of our approach for provably efficient IPE in broader applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。