arXiv:2604.24074cs.CL2026-04中稿 · the 22nd Internati…被引 2

安全评测中提示词设计会显著影响结果,导致模型安全评分波动达24.2个百分点。

How Sensitive Are Safety Benchmarks to Judge Configuration Choices?

论文配图:How Sensitive Are Safety Benchmarks to Judge Configuration Choices?
图 1 · 摘自论文原文
  • 通过12种提示变体测试,固定模型下提示词差异可使有害响应率变化24.2个百分点
  • 同一类别内微调提示词也能引发高达20.1个百分点的测量波动
  • 评测结果受提示设计和判断模型双重影响,适合评估评测系统稳定性的研究者参考

当前安全评测如HarmBench依赖大语言模型(LLM)作为评判者对模型输出进行有害性分类,但评判配置(即评判模型与提示词组合)通常被视为固定实现细节。本文采用2×2×3因子设计,构建12种提示变体,基于单一评判模型Claude Sonnet 4-6,在6个目标模型和400种HarmBench行为上生成28,812条判断。结果显示,仅提示词表述变化便使测得有害响应率最高波动24.2个百分点,同类内微调提示词亦能引发最大20.1个百分点的波动。模型安全排名中等不稳定,平均肯德尔tau为0.89;不同类别敏感度差异显著,版权类最高达39.6个百分点,而骚扰类为0。补充的多模型实验表明,评判模型选择进一步增加测量方差。结果表明,评判提示词是安全评测中此前被忽视的重要测量变异源。

原文摘要 · Abstract (English)

Safety benchmarks such as HarmBench rely on LLM judges to classify model responses as harmful or safe, yet the judge configuration, namely the combination of judge model and judge prompt, is typically treated as a fixed implementation detail. We show this assumption is problematic. Using a 2 x 2 x 3 factorial design, we construct 12 judge prompt variants along two axes, evaluation structure and instruction framing, and apply them using a single judge model, Claude Sonnet 4-6, producing 28,812 judgments over six target models and 400 HarmBench behaviors. We find that prompt wording alone, holding the judge model fixed, shifts measured harmful-response rates by up to 24.2 percentage points, with even within-condition surface rewording causing swings of up to 20.1 percentage points. Model safety rankings are moderately unstable, with mean Kendall tau = 0.89, and category-level sensitivity ranges from 39.6 percentage points for copyright to 0 percentage points for harassment. A supplementary multi-judge experiment using three judge models shows that judge-model choice adds further variance. Our results demonstrate that judge prompt wording is a substantial, previously under-examined source of measurement variance in safety benchmarking.

安全评测提示工程模型评估测评偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。