用配对比较优化设计,实现与传统方法相当的收敛速度。
A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
- 基于单调映射建模隐含效用差,结合双人核进行偏好反馈优化。
- 有限时间分析表明性能等同于标准贝叶斯优化的采样方法。
- 适合人类或专家参与的实验设计,尤其适用于无法获取数值评分场景。
偏好反馈(如成对比较而非数值评分)在人机协同设计、实验室实验及科学发现中日益重要。本文提出一种针对偏好反馈的汤普森采样(Thompson Sampling, TS)方法,通过单调链接函数建模潜在效用差异,并利用基础核函数诱导的双人核。有限时间分析显示,该方法性能与常规贝叶斯优化中标准TS相当。分析利用了TS在挑战者选择中的锚定不变性,并引入双TS配对变体。在合成数据和真实世界案例中均验证了其有效性。
原文摘要 · Abstract (English)
Preference feedback, in the form of pairwise comparisons rather than scalar scores, has seen increasing use in applications such as human-, laboratory-, and expert-in-the-loop design, as well as scientific discovery. We propose a Thompson Sampling (TS) approach to Bayesian optimization with preferential feedback that models comparisons using a monotone link on latent utility differences and leverages the dueling kernel induced by a base kernel. We provide a finite-time analysis showing that the performance of the proposed method matches that of standard TS for conventional Bayesian optimization with scalar feedback. The analysis exploits the anchor invariance of TS for challenger selection and introduces a double-TS pairing variant. We also demonstrate the performance of the method on both synthetic and real-world examples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。