arXiv:2409.05798cs.LGcs.AI2024-09NeurIPS被引 4

用用户反应时间提升偏好学习精度,加速找到最优选项

Enhancing Preference-based Linear Bandits via Human Response Time

  • 结合选择结果与反应时间,利用心理模型增强偏好估计
  • 在强偏好场景下,反应时间使效用估计误差降低40%以上
  • 适合需要快速决策的交互式推荐系统研究者

交互式偏好学习系统通过成对选项的二元选择来推断人类偏好。尽管二元选择简单易用,但仅能提供有限的偏好强度信息。为此,我们引入人类反应时间作为补充信号——其与偏好强度呈负相关。提出一种计算高效的联合方法,基于心理学中的EZ扩散模型,融合选择与反应时间以估计人类效用函数。理论与实证分析表明,在强偏好条件下,反应时间可补充选择信息,显著提升效用估计精度。将该估计器应用于固定预算的最佳臂识别任务中。在三个真实数据集上的模拟实验显示,相较于仅依赖选择的方法,使用反应时间可显著加速偏好学习过程。额外材料(代码、幻灯片、演讲视频)见https://shenlirobot.github.io/pages/NeurIPS24.html。

原文摘要 · Abstract (English)

Interactive preference learning systems infer human preferences by presenting queries as pairs of options and collecting binary choices. Although binary choices are simple and widely used, they provide limited information about preference strength. To address this, we leverage human response times, which are inversely related to preference strength, as an additional signal. We propose a computationally efficient method that combines choices and response times to estimate human utility functions, grounded in the EZ diffusion model from psychology. Theoretical and empirical analyses show that for queries with strong preferences, response times complement choices by providing extra information about preference strength, leading to significantly improved utility estimation. We incorporate this estimator into preference-based linear bandits for fixed-budget best-arm identification. Simulations on three real-world datasets demonstrate that using response times significantly accelerates preference learning compared to choice-only approaches. Additional materials, such as code, slides, and talk video, are available at https://shenlirobot.github.io/pages/NeurIPS24.html

偏好学习人机交互强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。