arXiv:2602.05110cs.AI2026-02

用多模型框架测评大模型在商户风险评估中的判别能力与偏见。

Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment

  • 构建五维度评分体系与蒙特卡洛打分,量化大模型推理质量与稳定性。
  • 发现部分模型存在显著自评偏差,匿名化使偏差降低25.8%。
  • 适合金融风控领域研究者及大模型应用落地团队参考。

大型语言模型(LLMs)被越来越多用于评估推理质量,但在支付风险场景下的可靠性和偏见尚不明确。本文提出一种结构化的多评估者框架,用于评估基于商户类别代码(MCC)的商户风险判断中的大模型推理表现,结合五准则评分标准与蒙特卡洛打分方法,对五个前沿大模型在有标识与匿名条件下生成并交叉评估风险理由。为建立独立于人类评判的基准,引入共识偏离度量,通过将每个评估者的得分与其余所有人的均值对比,消除循环评价问题,获得理论支撑的自我评估与跨模型偏差度量。结果表明存在显著异质性:GPT-5.1与Claude 4.5 Sonnet呈现负向自评偏差(-0.33、-0.31),而Gemini-2.5 Pro与Grok 4则显示正向偏差(+0.77、+0.71),匿名化使偏差减弱25.8%。26位支付行业专家评估显示,大模型评分平均高于人类共识0.46分;且GPT-5.1与Claude 4.5 Sonnet的负偏差反映出更接近人类判断。基于支付网络真实数据的验证表明,四个模型表现出统计显著的相关性(斯皮尔曼等级相关系数rho = 0.56至0.77),证实该框架能捕捉真实质量。整体而言,该框架为支付风险工作流中“大模型作为裁判”系统提供了可复现的评估基础,并强调了在金融操作场景中采用偏见感知协议的必要性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used as evaluators of reasoning quality, yet their reliability and bias in payments-risk settings remain poorly understood. We introduce a structured multi-evaluator framework for assessing LLM reasoning in Merchant Category Code (MCC)-based merchant risk assessment, combining a five-criterion rubric with Monte-Carlo scoring to evaluate rationale quality and evaluator stability. Five frontier LLMs generate and cross-evaluate MCC risk rationales under attributed and anonymized conditions. To establish a judge-independent reference, we introduce a consensus-deviation metric that eliminates circularity by comparing each judge's score to the mean of all other judges, yielding a theoretically grounded measure of self-evaluation and cross-model deviation. Results reveal substantial heterogeneity: GPT-5.1 and Claude 4.5 Sonnet show negative self-evaluation bias (-0.33, -0.31), while Gemini-2.5 Pro and Grok 4 display positive bias (+0.77, +0.71), with bias attenuating by 25.8 percent under anonymization. Evaluation by 26 payment-industry experts shows LLM judges assign scores averaging +0.46 points above human consensus, and that the negative bias of GPT-5.1 and Claude 4.5 Sonnet reflects closer alignment with human judgment. Ground-truth validation using payment-network data shows four models exhibit statistically significant alignment (Spearman rho = 0.56 to 0.77), confirming that the framework captures genuine quality. Overall, the framework provides a replicable basis for evaluating LLM-as-a-judge systems in payment-risk workflows and highlights the need for bias-aware protocols in operational financial settings.

大模型评估风险判断支付安全偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。