arXiv:2602.01600cs.CRcs.CL2026-02

提出预期危害度量,揭示大模型对低风险攻击防御过强、高风险攻击却易被突破的系统性缺陷。

Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs

  • 用执行成本建模攻击成功率,加权计算威胁的预期危害
  • 发现模型对高成本攻击拒绝率过高,对低成本攻击防护不足
  • 适合安全评估者和对抗攻击研究者参考

当前大模型安全评估多依赖基于严重性的分类体系,但该方法假设所有恶意请求风险均等,忽略了执行可能性——即模型响应后威胁实际发生的条件概率。本文提出预期危害指标,将攻击严重性与其执行可能性(由执行成本建模)结合。对前沿模型的实证分析显示存在系统性逆风险校准:模型对高成本(低概率)攻击拒绝强烈,却对低成本(高概率)攻击仍显脆弱。利用此特性,可使现有越狱攻击成功率最高提升2倍。通过线性探测进一步发现,模型在隐空间中能编码严重性以驱动拒绝行为,却无明确内部表示执行成本,导致对风险关键维度‘视而不见’。

原文摘要 · Abstract (English)

Current evaluations of LLM safety predominantly rely on severity-based taxonomies to assess the harmfulness of malicious queries. We argue that this formulation requires re-examination as it assumes uniform risk across all malicious queries, neglecting Execution Likelihood--the conditional probability of a threat being realized given the model's response. In this work, we introduce Expected Harm, a metric that weights the severity of a jailbreak by its execution likelihood, modeled as a function of execution cost. Through empirical analysis of state-of-the-art models, we reveal a systematic Inverse Risk Calibration: models disproportionately exhibit stronger refusal behaviors for low-likelihood (high-cost) threats while remaining vulnerable to high-likelihood (low-cost) queries. We demonstrate that this miscalibration creates a structural vulnerability: by exploiting this property, we increase the attack success rate of existing jailbreaks by up to $2\times$. Finally, we trace the root cause of this failure using linear probing, which reveals that while models encode severity in their latent space to drive refusal decisions, they possess no distinguishable internal representation of execution cost, making them "blind" to this critical dimension of risk.

大模型安全越狱攻击风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。