arXiv:2508.10390cs.CLcs.CR2025-08被引 1

用恶意提示设计新攻击,能突破最新大模型安全防护。

Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts

  • 整合多种攻击技巧生成伪造推理链,诱导模型输出有害内容。
  • 在GPT-5和Claude-4上成功率显著高于现有方法,提升超30%。
  • 构建新评估数据集RTA,有效剔除无效样本,更真实反映攻击效果。

现有黑盒越狱攻击在非推理模型上有效,但在最新SOTA推理模型上表现下降。受对抗聚合策略启发,本文将多种越狱技巧集成到统一开发模板中,引入对抗上下文对齐消除语义矛盾,并使用NTP(有害提示)作为少样本示例引导恶意输出,最终形成带有虚假推理链的DH-CoT攻击。实验发现,现有红队测试数据集包含不适用于评估攻击效果的样本,如BPs、NHPs和NTPs,干扰真实攻击效果评估。为此,提出MDH框架,结合大模型标注与人工辅助实现恶意内容检测,清理数据并构建RTA数据集套件。结果表明,MDH能可靠过滤低质量样本,且DH-CoT可有效越狱GPT-5和Claude-4,显著优于当前最优方法H-CoT和TAP。

原文摘要 · Abstract (English)

Existing black-box jailbreak attacks achieve certain success on non-reasoning models but degrade significantly on recent SOTA reasoning models. To improve attack ability, inspired by adversarial aggregation strategies, we integrate multiple jailbreak tricks into a single developer template. Especially, we apply Adversarial Context Alignment to purge semantic inconsistencies and use NTP (a type of harmful prompt) -based few-shot examples to guide malicious outputs, lastly forming DH-CoT attack with a fake chain of thought. In experiments, we further observe that existing red-teaming datasets include samples unsuitable for evaluating attack gains, such as BPs, NHPs, and NTPs. Such data hinders accurate evaluation of true attack effect lifts. To address this, we introduce MDH, a Malicious content Detection framework integrating LLM-based annotation with Human assistance, with which we clean data and build RTA dataset suite. Experiments show that MDH reliably filters low-quality samples and that DH-CoT effectively jailbreaks models including GPT-5 and Claude-4, notably outperforming SOTA methods like H-CoT and TAP.

越狱攻击大模型安全对抗样本红队测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。