arXiv:2601.09084cs.CLcs.LG2026-01

小改进难被发现,人类评价需更多判断才能可靠检测。

How Many Human Judgments Are Enough? Feasibility Limits of Human Preference Evaluation

  • 当提示信息分散时,按比例分配判断最有效。
  • 多数评估中偏好差距小,需远超常规数量的判断才能发现改进。
  • 精心设计的基准能降低提示差异,提升检测能力1.5倍。

人类偏好评估广泛用于对比生成模型,但尚不清楚需要多少判断才能可靠检测微小改进。我们发现,当偏好信号在各类提示间均匀分布(即所有提示类型都具有相似信息量)时,按比例分配判断为最小最大最优策略:任何分配方式都无法显著提升检测能力。对大规模人类偏好数据集的实证分析显示,大多数比较处于此分散状态,偏好差距微小,所需判断数量远超通常收集量,即使在充分采样的比较中也是如此。这些限制在不同评估协议和模态下均存在,涵盖对话、图像生成及带执行反馈的代码生成。相反,通过筛选提示减少提示引起的变异性,可系统性增大偏好差距,使提示层面方差降低1.5倍,显著提升检测能力。结果表明,评估结果不显著或为负值常源于统计功效不足,而非模型无差异,强调必须显式考虑效应大小、预算与评估设计。

原文摘要 · Abstract (English)

Human preference evaluations are widely used to compare generative models, yet it remains unclear how many judgments are required to reliably detect small improvements. We show that when preference signal is diffuse across prompts (i.e., all prompt types are similarly informative), proportional allocation is minimax-optimal: no allocation strategy substantially improves detectability. Empirical analysis of large-scale human preference datasets shows that most comparisons fall into this diffuse regime, exhibiting small preference margins that require far more judgments than typically collected, even in well-sampled comparisons. These limits persist across evaluation protocols and modalities, including chat, image generation, and code generation with execution feedback. In contrast, curated benchmarks that reduce prompt induced variability systematically induce larger margins and improve detectability through a $1.5\times$ reduction in prompt-level variance. Our results show that inconclusive or negative human evaluation outcomes frequently reflect underpowered evaluation rather than model equivalence, underscoring the need to account explicitly for effect size, budget, and protocol design.

人类评估生成模型统计功效提示设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。