arXiv:2509.23882cs.AIcs.CR2025-09

探查GPT-OSS-20B在对抗攻击下的五类失效模式,揭示安全漏洞。

Quant Fever, Reasoning Blackholes, Schrodinger's Compliance, and More: Probing GPT-OSS-20B

  • 用系统化工具探测模型在不同攻击下的行为异常。
  • 发现量化过热、推理黑洞等五类严重失效现象。
  • 适合关注大模型安全与对抗鲁棒性的研究者阅读。

OpenAI的GPT-OSS系列提供具有显式思维链(CoT)推理和和谐提示格式的开源权重语言模型。本文对GPT-OSS-20B进行了全面的安全评估,通过Jailbreak Oracle(JO)这一系统化LLM评估工具,探测模型在各类对抗条件下的行为表现。实验揭示了包括量化过热(quant fever)、推理黑洞(reasoning blackholes)、薛定谔合规(Schrodinger's compliance)、推理过程幻觉(reasoning procedure mirage)以及链式提示依赖(chain-oriented prompting)在内的多种失效模式。这些行为可被恶意利用,导致严重后果,凸显了当前开源大模型在安全性方面的潜在风险。

原文摘要 · Abstract (English)

OpenAI's GPT-OSS family provides open-weight language models with explicit chain-of-thought (CoT) reasoning and a Harmony prompt format. We summarize an extensive security evaluation of GPT-OSS-20B that probes the model's behavior under different adversarial conditions. Using the Jailbreak Oracle (JO) [1], a systematic LLM evaluation tool, the study uncovers several failure modes including quant fever, reasoning blackholes, Schrodinger's compliance, reasoning procedure mirage, and chain-oriented prompting. Experiments demonstrate how these behaviors can be exploited on the GPT-OSS-20B model, leading to severe consequences.

大模型安全对抗攻击思维链漏洞挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。