arXiv:2510.06538cs.AIcs.LG2025-10被引 4

让大模型评委自动补全评价维度,提升判断可靠性

Auto-Prompt Ensemble for LLM Judge

  • 通过自动学习失败案例中的新评价维度来增强模型判断
  • 零样本下使GPT-4o在Reward Bench的共识率从87.2%提升至90.5%
  • 适合需要高可信度评估的AI评测场景

我们提出一种新框架Auto-Prompt Ensemble(APE),通过选择性地为大模型评委补充辅助评价维度,提升其判断可靠性。现有大模型评委常因未能识别人类评估背后的隐含标准而遗漏关键维度。APE采用自适应机制,从失败案例中自动学习新评价维度,并引入基于置信度的集成方法——集体置信度(Collective Confidence),决定何时采纳额外维度的判断。大量实验表明,APE在多个标准基准上显著提升了大模型评委的可靠性。例如,在零样本设置下,APE将GPT-4o在Reward Bench上的共识率从87.2%提升至90.5%。整体而言,APE为大模型评委提供了利用测试时计算的系统性方法,有效缩小了人类与大模型评委之间的评估差距。

原文摘要 · Abstract (English)

We present a novel framework that improves the reliability of LLM judges by selectively augmenting LLM with auxiliary evaluation dimensions. Existing LLM judges often miss crucial evaluation dimensions because they fail to recognize the implicit standards underlying human assessments. To address this challenge, we propose the Auto-Prompt Ensemble (APE), an adaptive framework that automatically learns evaluation dimensions from its failure cases. APE incorporates a confidence-based ensemble mechanism to decide when to adopt the judgments from additional evaluation dimensions through a novel confidence estimation approach called Collective Confidence. Extensive experiments demonstrate that APE improves the reliability of LLM Judge across diverse standard benchmarks. For instance, APE enhances GPT-4o agreement rate on Reward Bench from 87.2% to 90.5% in the zero-shot setting. Overall, APE provides a principled approach for LLM Judge to leverage test-time computation, and bridge the evaluation gap between human and LLM judges.

大模型评测自动提示置信度集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。