从少量比较数据中学习复杂分布,无需直接标注。
Score-Based Density Estimation from Pairwise Comparisons
- 通过比较结果反推偏好得分,再还原真实分布
- 仅用数百到数千次比较即可估计多维复杂密度
- 适合人类反馈或专家知识建模的场景
我们研究基于成对比较的密度估计,动机来自专家知识获取和从人类反馈中学习。将未观测的目标密度与经过温度调整的胜者密度(优选选择的边际密度)关联,通过得分匹配学习胜者得分,进而通过去温度化估计目标密度。我们证明信念密度与胜者密度的得分向量共线,由位置相关的温度场连接。给出了该场的解析表达式,并在Bradley-Terry模型下提出估计方法。利用在温度化样本上训练的扩散模型,这些样本通过得分加权的退火Langevin动力学生成,我们仅需数百至数千次成对比较,即可学习模拟专家的复杂多变量信念密度。
原文摘要 · Abstract (English)
We study density estimation from pairwise comparisons, motivated by expert knowledge elicitation and learning from human feedback. We relate the unobserved target density to a tempered winner density (marginal density of preferred choices), learning the winner's score via score-matching. This allows estimating the target by `de-tempering' the estimated winner density's score. We prove that the score vectors of the belief and the winner density are collinear, linked by a position-dependent tempering field. We give analytical formulas for this field and propose an estimator for it under the Bradley-Terry model. Using a diffusion model trained on tempered samples generated via score-scaled annealed Langevin dynamics, we can learn complex multivariate belief densities of simulated experts, from only hundreds to thousands of pairwise comparisons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。