用温度参数动态调节AI评估严格度,适配不同应用场景。
Adaptive Rigor in AI System Evaluation using Temperature-Controlled Verdict Aggregation via Generalized Power Mean
- 引入温度控制的广义幂平均聚合,灵活调节评估严格程度。
- 在三个数据集上与人工评分相关性达0.667,优于DeepEval。
- 调参无需额外大模型调用,适合安全与对话等不同场景。
现有基于大模型的AI系统评估方法(如LLM-as-a-Judge、 verdict系统和NLI)常因无法适应应用领域而与人类判断不一致。本文提出温度控制的裁决聚合(TCVA),结合五级评分体系与广义幂平均聚合,并引入直观温度参数T∈[0.1,1.0]以控制评估严格度:低温度生成偏保守分数,适用于安全关键领域;高温度产生宽松分数,适合对话类AI。在包含人类Likert量表标注的三个基准数据集(SummEval和USR)上的实验表明,TCVA在忠实性方面与人工评分的相关性达到Spearman=0.667,接近RAGAS的0.676,且始终优于DeepEval。该方法调整温度时无需额外大模型调用。
原文摘要 · Abstract (English)
Existing evaluation methods for LLM-based AI systems, such as LLM-as-a-Judge, verdict systems, and NLI, do not always align well with human assessment because they cannot adapt their strictness to the application domain. This paper presents Temperature-Controlled Verdict Aggregation (TCVA), a method that combines a five-level verdict-scoring system with generalized power-mean aggregation and an intuitive temperature parameter T [0.1, 1.0] to control evaluation rigor. Low temperatures yield pessimistic scores suited for safety-critical domains; high temperatures produce lenient scores appropriate for conversational AI. Experimental evaluation on three benchmark datasets with human Likert-scale annotations (SummEval and USR) shows that TCVA achieves correlation with human judgments comparable to RAGAS on faithfulness (Spearman = 0.667 vs. 0.676) while consistently outperforming DeepEval. The method requires no additional LLM calls when adjusting the temperature parameter.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。