arXiv:2604.03742cs.AI2026-04

用模糊层次分析法让大模型评估更稳定可靠。

Structured Multi-Criteria Evaluation of Large Language Models with Fuzzy Analytic Hierarchy Process and DualJudge

  • 引入置信度调节的模糊层次分析,量化评估中的不确定性
  • 在JudgeBench上验证,比直接打分更稳定且准确
  • 融合直觉与系统评估,适合需要高可信度的评测场景

大语言模型(LLM)的有效评估仍是关键瓶颈,传统直接评分常导致结果不一致且难以解释。本文将层次分析法(AHP)应用于基于LLM的评估,并提出一种置信度感知的模糊层次分析(FAHP)扩展方法,通过三角模糊数建模认知不确定性,其参数由LLM生成的置信度动态调节。在JudgeBench上的系统性验证表明,该结构化方法将评估分解为明确标准并引入不确定性感知聚合,获得更校准的判断。大量实验显示,无论是清晰还是模糊的AHP,均在不同模型规模和数据划分下持续优于直接评分,其中FAHP在不确定对比场景中表现出更强稳定性。基于此,我们提出双判官(DualJudge)框架,受双过程理论启发,通过一致性感知加权自适应融合整体直觉评分与结构化AHP输出。DualJudge达到当前最优性能,凸显直觉与深思评估范式的互补优势。这些结果确立了不确定性感知的结构化推理作为更可靠大模型评估的可遵循路径。代码已开源:https://github.com/hreyulog/AHP_llm_judge。

原文摘要 · Abstract (English)

Effective evaluation of large language models (LLMs) remains a critical bottleneck, as conventional direct scoring often yields inconsistent and opaque judgments. In this work, we adapt the Analytic Hierarchy Process (AHP) to LLM-based evaluation and, more importantly, propose a confidence-aware Fuzzy AHP (FAHP) extension that models epistemic uncertainty via triangular fuzzy numbers modulated by LLM-generated confidence scores. Systematically validated on JudgeBench, our structured approach decomposes assessments into explicit criteria and incorporates uncertainty-aware aggregation, producing more calibrated judgments. Extensive experiments demonstrate that both crisp and fuzzy AHP consistently outperform direct scoring across model scales and dataset splits, with FAHP showing superior stability in uncertain comparison scenarios. Building on these insights, we propose \textbf{DualJudge}, a hybrid framework inspired by Dual-Process Theory that adaptively fuses holistic direct scores with structured AHP outputs via consistency-aware weighting. DualJudge achieves state-of-the-art performance, underscoring the complementary strengths of intuitive and deliberative evaluation paradigms. These results establish uncertainty-aware structured reasoning as a principled pathway toward more reliable LLM assessment. Code is available at https://github.com/hreyulog/AHP_llm_judge.

大模型评估模糊逻辑双过程理论可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。