用偏好问题精准提取专家高维概率分布,突破传统方法局限。
Preferential Normalizing Flows

- 基于偏好问答构建归一化流模型,无需直接评分
- 提出新函数先验有效防止概率质量坍缩或发散
- 适用于大模型先验建模等需要灵活密度估计场景
通过噪声判断从专家处获取高维概率分布极具挑战性,但在先验设定和奖励建模等任务中极为有用。本文提出一种仅依赖偏好问题(如比较或排序)来构建归一化流模型的方法,从而原则上可建模任意复杂的分布。然而,流模型估计易受概率质量坍缩或发散的影响。为此,我们引入一种基于决策理论动机的新函数先验,并实证表明可将专家信念密度作为函数空间的极大后验估计。我们在模拟专家上验证了多变量信念密度的提取能力,包括通用大语言模型在真实数据集上的先验信念建模。
原文摘要 · Abstract (English)
Eliciting a high-dimensional probability distribution from an expert via noisy judgments is notoriously challenging, yet useful for many applications, such as prior elicitation and reward modeling. We introduce a method for eliciting the expert's belief density as a normalizing flow based solely on preferential questions such as comparing or ranking alternatives. This allows eliciting in principle arbitrarily flexible densities, but flow estimation is susceptible to the challenge of collapsing or diverging probability mass that makes it difficult in practice. We tackle this problem by introducing a novel functional prior for the flow, motivated by a decision-theoretic argument, and show empirically that the belief density can be inferred as the function-space maximum a posteriori estimate. We demonstrate our method by eliciting multivariate belief densities of simulated experts, including the prior belief of a general-purpose large language model over a real-world dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。