arXiv:2509.14297cs.CRcs.CL2025-09被引 4

用学习类提问骗大模型输出有害内容,成功率高且难防御。

A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness

  • 通过重构问题结构,伪装成学习提问诱导模型越界。
  • 在多个模型上攻击成功率超90%,提示词极简高效。
  • 揭示现有安全机制缺陷,适合安全研究者参考。

本研究揭示现代大模型存在关键安全盲点:以学习风格重构的查询(如模仿教育问答)能稳定诱导出有害响应。提出新型重述框架HILL(Hiding Intention by Learning from LLMs),包含四大组件:核心概念、探索性转换、细节追问,以及可选假设性表述。该方法为确定性、模型无关的重述策略,在广泛模型上的AdvBench数据集测试中表现出强泛化能力,多数模型与恶意类别下攻击成功率领先,且提示词简洁高效。同时,多种防御手段测试显示,多数防御效果平庸甚至提升攻击成功率。对安全提示的评估进一步暴露大模型安全机制的内在局限及防御方法的缺陷。本工作凸显了在保持模型助人属性的同时实现安全对齐的重大挑战。

原文摘要 · Abstract (English)

This study reveals a critical safety blind spot in modern LLMs: learning-style queries, which closely resemble ordinary educational questions, can reliably elicit harmful responses. The learning-style queries are constructed by a novel reframing paradigm: HILL (Hiding Intention by Learning from LLMs). The deterministic, model-agnostic reframing framework is composed of 4 conceptual components: 1) key concept, 2) exploratory transformation, 3) detail-oriented inquiry, and optionally 4) hypotheticality. Further, new metrics are introduced to thoroughly evaluate the efficiency and harmfulness of jailbreak methods. Experiments on the AdvBench dataset across a wide range of models demonstrate HILL's strong generalizability. It achieves top attack success rates on the majority of models and across malicious categories while maintaining high efficiency with concise prompts. On the other hand, results of various defense methods show the robustness of HILL, with most defenses having mediocre effects or even increasing the attack success rates. In addition, the assessment of defenses on the constructed safe prompts reveals inherent limitations of LLMs' safety mechanisms and flaws in the defense methods. This work exposes significant vulnerabilities of safety measures against learning-style elicitation, highlighting a critical challenge of fulfilling both helpfulness and safety alignments.

模型安全越狱攻击提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。