arXiv:2603.20562cs.CLcs.AI2026-03中稿 · the Fifth Workshop…被引 1

通过多次打乱答案顺序提升大模型事实性评估的稳定性

Permutation-Consensus Listwise Judging for Robust Factuality Evaluation

  • 对同一组答案多次打乱顺序后统一评分,生成共识结果
  • 在RewardBench 2数据集上将准确率从86%提升至91.33%
  • 适合需要高可靠性的模型评估场景,尤其关注幻觉检测

大语言模型现常被用作评判者,但其判断易受无关呈现方式影响。本文研究列表式事实性评估中的候选答案顺序敏感性问题——多个答案外观相似,但幻觉风险差异显著。提出PCFJudge方法,在推理阶段对同一候选集多次使用不同排序的列表式提示,并聚合得分、排名与不确定性信号,形成共识决策。在RewardBench 2 Factuality测试中,采用7种排列组合的最终平均结果,使GPT-5.4的顶1选择准确率从86.00%提升至91.33%,Claude Sonnet 4.6从86.33%提升至89.67%。结果表明,答案顺序是事实性评判误差的重要来源,通过边缘化该干扰因素可显著提升评估可靠性。

原文摘要 · Abstract (English)

Large language models (LLMs) are now widely used as judges, yet their decisions can change under presentation choices that should be irrelevant. We study one such source of instability: candidate-order sensitivity in listwise factuality evaluation, where several answers can look similarly polished while differing substantially in hallucination risk. We introduce PCFJudge, an inference-time method that reruns the same factuality-first listwise prompt over multiple orderings of the same candidate set and aggregates the resulting scores, ranks, and uncertainty signals into a single consensus decision. On RewardBench 2 Factuality, the final seven-permutation aggregate (K=7) improves top-1 selection accuracy from 86.00% to 91.33% with GPT-5.4 and from 86.33% to 89.67% with Claude Sonnet 4.6. These results suggest that candidate order can be a meaningful source of factuality-judging error and that marginalizing over this nuisance variation can improve the reliability of LLM evaluation.

事实性评估大模型评测鲁棒性共识决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。