arXiv:2510.09333cs.LGcs.CV2025-10被引 2

用贝叶斯方法改进人工评分,自动识别并过滤低质量评价者。

Efficient Bayesian Inference from Noisy Pairwise Comparisons

  • 引入贝叶斯布朗-泰瑞模型,显式建模评分者质量
  • 在嘈杂或众包数据下仍保持排名稳定与不确定性校准
  • 适合需要低成本高可靠人类评估的生成模型研究

评估生成模型困难,因标准指标常无法反映人类偏好。人工评估更可靠但成本高且含噪声,参与者专业度、注意力和投入程度不一。成对比较可提升一致性,但聚合为总体质量评分需精细建模。基于布朗-泰瑞的方法从比较中更新项目得分,但现有方法或忽略评分者差异,或缺乏收敛性保证,影响鲁棒性与可解释性。本文提出BBQ,一种显式建模评分者质量的贝叶斯布朗-泰瑞变体,通过期望最大化算法实现单调似然收敛,并自动降低或移除不可靠评分者的影响。实验证明,即便在噪声或众包评分条件下,BBQ仍能实现高效推断、校准的不确定性估计及更稳健、可解释的排名,优于基线布朗-泰瑞模型。该框架使生成模型的人类评估更可靠且成本更低。

原文摘要 · Abstract (English)

Evaluating generative models is challenging because standard metrics often fail to reflect human preferences. Human evaluations are more reliable but costly and noisy, as participants vary in expertise, attention, and diligence. Pairwise comparisons improve consistency, yet aggregating them into overall quality scores requires careful modeling. Bradley-Terry-based methods update item scores from comparisons, but existing approaches either ignore rater variability or lack convergence guarantees, limiting robustness and interpretability. We introduce BBQ, a Bayesian Bradley-Terry variant that explicitly models rater quality, downweighting or removing unreliable participants, and provides guaranteed monotonic likelihood convergence through an Expectation-Maximization algorithm. Empirical results show that BBQ provides efficient inference, well-calibrated uncertainty estimates, and more robust, interpretable rankings compared to baseline Bradley-Terry models, even with noisy or crowdsourced raters. This framework enables more reliable and cost-effective human evaluation of generative models.

人类评估贝叶斯推理生成模型评分建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。