arXiv:2512.03019cs.LGcs.AI2025-12被引 2

通过校准分布的计算分配,让大模型评分更可靠。

Distribution-Calibrated Inference Time Compute for Thinking LLM-as-a-Judge

  • 用计数建模三元偏好,结合极性与确定性判断
  • 在多个基准上降低平均绝对误差,提升配对准确率
  • 适合需要高可靠性评估的场景,如评测模型质量

以思维链大模型作为评判者进行成对偏好判断时,单次样本仍存在噪声,常见聚合规则(多数投票、软自一致性或指令自聚合)在允许平局时表现不一致。本文研究为每个项目生成n个独立思维-评分样本的推理时计算(ITC),提出一种基于分布校准的聚合方法。该方法采用Bradley-Terry-Davidson模型对评分计数建模,同时利用极性(非平局间的差距)和果断性(非平局率)区分微弱差距与强共识。在多个评估基准上,本方法持续降低平均绝对误差(MAE),提升配对准确率;与人类共识元标签对比,表现匹配甚至超过个别真人评判者。结果表明,合理分配推理时计算并使用分布感知聚合,可将噪声个体判断转化为可靠的评估评分。

原文摘要 · Abstract (English)

Thinking Large Language Models (LLMs) used as judges for pairwise preferences remain noisy at the single-sample level, and common aggregation rules (majority vote, soft self-consistency, or instruction-based self-aggregation) are inconsistent when ties are allowed. We study inference-time compute (ITC) for evaluators that generate n independent thinking--rating samples per item, and propose a principled, distribution-calibrated aggregation scheme. Our method models three-way preferences with a Bradley-Terry-Davidson formulation on rating counts, leveraging both polarity (margin among non-ties) and decisiveness (non-tie rate) to distinguish narrow margins from strong consensus. Across various evaluation benchmarks, our approach consistently reduces MAE and increases pairwise accuracy versus standard baselines, and when evaluated against human-consensus meta-labels, matches or exceeds individual human raters. These results show that carefully allocating ITC and aggregating with distribution-aware methods turns noisy individual model judgments into reliable ratings for evaluation.

大模型评测推理计算偏好判断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。