arXiv:2607.28641cs.CLcs.AI2026-07

大模型裁判易受群体思维误导,导致判断失真。

The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?

论文配图:The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?
图 1 · 摘自论文原文
  • 提出评估者失调指数,量化模型对形式化结构的盲从
  • 在2.25万条数据中发现幻觉行为的语义分类体系
  • 证明漏洞普遍存在且与领域无关,需针对性防护

我们提出「代理形式主义陷阱」和评估不一致指数(D_E),量化大型语言模型作为裁判时,在对抗性压力下将形式程序性误认为语义真实性的现象。分析了跨3个领域的22,500条推理轨迹(GAIA、SWE-bench、Multi-Challenge),提取出经确定性词法锚定验证的幻觉操作语义分类体系(p < 10^{-120})。逻辑元评估器识别出触发该评估捕获的精确句法特征(ROC-AUC 0.8779),零样本留一域外迁移测试表明该漏洞具有普遍性(平均ROC-AUC 0.7482)。架构剖析显示,不同模拟群体拓扑结构引发数学上不同的语义盲区,证明无锚点闭环评估系统不稳定、系统性发散,必须采用架构定制的警戒过滤机制。

原文摘要 · Abstract (English)

We introduce the \textit{Agentic Formalism Trap} and the Evaluative Dissonance Index ($D_E$), quantifying how LLM-as-a-Judge systems conflate structural proceduralism with semantic truth under adversarial load. Analyzing 22,500 trajectories across 3 domains (GAIA, SWE-bench, Multi-Challenge), we extract a semantic taxonomy of hallucination maneuvers, validated via deterministic lexical grounding ($p < 10^{-120}$). A logistic meta-evaluator isolates the exact syntactic triggers of this evaluator capture (ROC-AUC 0.8779), while a zero-shot Leave-One-Domain-Out transfer proves the vulnerability is universally domain-agnostic (mean ROC-AUC 0.7482). Architectural profiling reveals that distinct simulated swarm topologies induce mathematically disparate semantic blind spots, proving that unanchored closed-loop evaluation is unstable, systemically divergent and necessitates architecture-specific vigilance filters.

大模型评估幻觉检测形式主义陷阱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。