arXiv:2603.29403cs.CRcs.AI2026-03被引 3

首次系统梳理大模型作为裁判时的安全风险与防御策略。

Security in LLM-as-a-Judge: A Comprehensive SoK

  • 构建四类安全角色分类体系,厘清攻击与防御路径。
  • 分析45篇相关研究,揭示评估框架的显著漏洞。
  • 适合关注AI评测安全的科研人员与开发者参考。

LLM-as-a-Judge(LaaJ)是一种新型范式,利用强大语言模型评估生成输出的质量、安全性和正确性。尽管该范式大幅提升了评估的可扩展性与效率,但也引入了尚未充分探索的安全风险与可靠性问题。具体而言,基于大模型的裁判可能成为对抗性攻击的目标,或被用作实施攻击的工具,从而危及评估流程的可信度。本文首次针对LaaJ系统的安全性进行系统知识归纳(SoK),通过跨主要学术数据库的全面文献调研,分析863篇论文并筛选出2020至2026年间45篇相关研究。基于此,我们提出一个分类体系,按大模型裁判在安全场景中的角色,区分为:针对LaaJ的攻击、通过LaaJ实施的攻击、利用LaaJ进行防御、以及在安全相关领域中使用LaaJ作为评估策略的应用。进一步开展对比分析,指出现有方法的局限、新兴威胁与开放挑战。研究发现,大模型评估框架存在显著脆弱性,但亦展现出提升鲁棒性与可靠性的潜力。最后,我们提出了若干关键研究方向,以指导更安全、可信的LaaJ系统发展。

原文摘要 · Abstract (English)

LLM-as-a-Judge (LaaJ) is a novel paradigm in which powerful language models are used to assess the quality, safety, or correctness of generated outputs. While this paradigm has significantly improved the scalability and efficiency of evaluation processes, it also introduces novel security risks and reliability concerns that remain largely unexplored. In particular, LLM-based judges can become both targets of adversarial manipulation and instruments through which attacks are conducted, potentially compromising the trustworthiness of evaluation pipelines. In this paper, we present the first Systematization of Knowledge (SoK) focusing on the security aspects of LLM-as-a-Judge systems. We perform a comprehensive literature review across major academic databases, analyzing 863 works and selecting 45 relevant studies published between 2020 and 2026. Based on this study, we propose a taxonomy that organizes recent research according to the role played by LLM-as-a-Judge in the security landscape, distinguishing between attacks targeting LaaJ systems, attacks performed through LaaJ, defenses leveraging LaaJ for security purposes, and applications where LaaJ is used as an evaluation strategy in security-related domains. We further provide a comparative analysis of existing approaches, highlighting current limitations, emerging threats, and open research challenges. Our findings reveal significant vulnerabilities in LLM-based evaluation frameworks, as well as promising directions for improving their robustness and reliability. Finally, we outline key research opportunities that can guide the development of more secure and trustworthy LLM-as-a-Judge systems.

大模型评测安全风险系统综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。