同一提示的单一标准形式会低估大模型安全性的表面敏感性。
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity
- 用多种语义不变的改写形式测试模型安全响应,避免标准形式偏差。
- 单一标准形式下安全评估遗漏3.3%-12.9%的危险行为,部分原本安全的提示在改写后变危险。
- 结果表明模型对提示表面形式高度敏感,适合关注模型安全评估可靠性的研究者。
基准分数是测量工具,但多数基准仅以单一标准表面形式读取每项内容。我们质疑这种读取是否准确:当意图保持不变而表面形式变化时,标准形式分数能否真实反映模型行为?其中多少变异来自解码或评判噪声而非真实信号?我们在高风险的安全场景中进行验证,该场景无黄金标签可平均。为避免先验混淆,预先作者化改写(无拒绝、主要非LLM:机器回译与矩阵语言框架混杂生成器),使相同表面形式送达所有模型;使用一人锚定、厂商中立的裁判(Claude,与人类对比κ=0.86,跨语言稳定,经GPT-4o交叉验证)评分,并验证意图一致性。在370个种子×5种表面形式×5个模型下,无任一改写形式始终最危险(每形式经校正后仅6/20个McNemar检验显著,多数具保护性)。仅使用标准提示会低估危险合规率:各形式下危险结果的并集超过最差单一形式3.3-12.9个百分点,所有五模型的置信区间均不包含零;5-13%的种子在标准提示下安全,但在某些改写下变为危险——高于零随机性基线(标准提示温度0重采样五次得0/370新增暴露)。该差距大小因模型而异(最大于Gemini 2.5 Pro)。一种改写形式仅恢复约53%的模型可观测危险表面,约三者达85%——此为形式集合的冗余特征,非特定群体定义。良性对照(XSTest)表明不稳定性具有双向性,但良性与有害样本未逐项匹配。数据集、代码与逐响应标签已公开。
原文摘要 · Abstract (English)
A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is faithful: when an item's intent is held fixed and only its meaning-preserving surface form varies, does the canonical-form score estimate model behavior well, and how much of any variation is decoding/judge noise rather than signal? We instantiate this in safety, a high-stakes setting with no gold label to average toward. To avoid prior confounds, we pre-author the reformulations (refusal-free, mostly non-LLM: machine back-translation and a Matrix-Language-Frame code-switch generator) so an identical surface form reaches every model, score all responses with one human-anchored, vendor-neutral judge (Claude, kappa = 0.86 vs. human on unsafe compliance, stable across languages, cross-checked by GPT-4o), and verify intent preservation. On 370 seeds x 5 surface forms x 5 models, no single transformation is uniformly most dangerous (6 of 20 per-transformation McNemar tests survive correction, most protective). Yet evaluating only the canonical prompt underestimates unsafe compliance: the union of unsafe outcomes across forms exceeds even the worst single form by 3.3-12.9 pp, with bootstrap 95% CIs excluding zero for all five models, and 5-13% of seeds safe on canonical are unsafe under some reformulation -- above a zero stochasticity floor (canonical resampled five times at temperature 0 gives 0/370 new exposures). The size of this gap is model-dependent (largest on Gemini 2.5 Pro). One form recovers only ~53% of a model's observed unsafe surface and about three reach 85% -- a redundancy characterization of this form set, not of a defined population. A benign control (XSTest) suggests the instability is bidirectional, though the benign and harmful pools are not item-matched. We release the dataset, code, and per-response labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。