arXiv:2606.07874cs.AI2026-06被引 2

大模型判官易受上下文影响,安全判断不靠谱。

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators

论文配图:Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators
图 1 · 摘自论文原文
  • 用上下文信息和不同安全定义测试大模型判官
  • 多数模型不因新信息或定义改变原有安全判断
  • 适合关注评估可靠性的研究人员

以大语言模型作为评判者是大规模评估安全性的唯一途径。尽管其重要性突出,但现有研究多仅通过人类一致性在简单静态基准上评估其表现。我们深入探讨了两类未被充分研究的关键属性:大模型判官对上下文信息的依赖性,以及其对不同安全定义的可引导性,这些可能与其内部安全先验不一致。我们评估了多种通用大模型与专用安全判官的安全判断能力,并考察任务示例、新上下文信息及安全定义变化的影响。结果表明,尽管大模型能学习新信息,但当上下文或安全定义与其既有先验冲突时,普遍不会调整评价结果。

原文摘要 · Abstract (English)

LLMs-as-judges are the only way to evaluate safety at scale. Despite their importance, LLM-judges themselves are rarely evaluated beyond human agreement in simple, static benchmarks. We therefore investigate two under-explored but crucial properties of LLMs-as-judges: their susceptibility to relying on in context-information, and their steerability to differing safety definitions, which may not align with their internal safety priors. We evaluate the safety judging abilities of many generalist LLMs and safety-specific judges, and investigate the impact of task demonstrations, novel in-context information, and changing safety definitions. We find that while LLM-judges can learn from new information, they are broadly unlikely to adjust their evaluations if the context or safety definition contradicts their prior.

安全评估大模型判官上下文依赖先验偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。