arXiv:2510.08120cs.CLcs.AI2025-10被引 3

解析大模型评判文本的决策逻辑,揭示其隐藏规则与潜在偏见。

Interpreting LLM-as-a-Judge Policies via Verifiable Global Explanations

  • 通过对比解释与聚类归纳,从大模型评判中提炼出可验证的全局规则。
  • 在7个基准数据集上验证,提取规则与模型决策高度一致。
  • 适用于安全审核、内容风控等需透明化评估的场景。

利用大模型进行文本评价(即LLM-as-a-judge)正被大规模采用以补充甚至替代人工标注。因此,理解其潜在偏见与风险至关重要。本文提出一种方法,从LLM-as-a-judge中提取基于概念的全局政策。该方法包含两个算法:1)CLoVE(对比局部可验证解释),生成基于概念、可验证的对比性局部解释;2)GloVE(全局可验证解释),通过迭代聚类、摘要与验证,将局部规则压缩为全局政策。我们在七个标准内容危害检测数据集上评估了GloVE,结果表明提取的全局政策与大模型决策高度忠实。此外,我们评估了全局政策对文本扰动和对抗攻击的鲁棒性,并通过用户研究验证了用户对全局政策的理解度与满意度。

原文摘要 · Abstract (English)

Using LLMs to evaluate text, that is, LLM-as-a-judge, is increasingly being used at scale to augment or even replace human annotations. As such, it is imperative that we understand the potential biases and risks of doing so. In this work, we propose an approach for extracting high-level concept-based global policies from LLM-as-a-Judge. Our approach consists of two algorithms: 1) CLoVE (Contrastive Local Verifiable Explanations), which generates verifiable, concept-based, contrastive local explanations and 2) GloVE (Global Verifiable Explanations), which uses iterative clustering, summarization and verification to condense local rules into a global policy. We evaluate GloVE on seven standard benchmarking datasets for content harm detection. We find that the extracted global policies are highly faithful to decisions of the LLM-as-a-Judge. Additionally, we evaluated the robustness of global policies to text perturbations and adversarial attacks. Finally, we conducted a user study to evaluate user understanding and satisfaction with global policies.

大模型评估可解释性内容安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。