arXiv:2507.08014cs.CLcs.AI2025-07

分析两百万次真实对话,发现大模型越狱攻击复杂度并未持续升级。

Mass-Scale Analysis of In-the-Wild Conversations Reveals Complexity Bounds on LLM Jailbreaking

  • 通过多种复杂度指标分析真实对话,识别越狱策略特征。
  • 越狱攻击复杂度与正常对话相当,无明显上升趋势。
  • 适合关注AI安全、对抗攻击的开发者与研究者阅读。

随着大语言模型(LLMs)广泛应用,理解越狱策略的复杂性与演化至关重要。我们对来自不同平台的超过200万条真实对话进行了大规模实证分析,涵盖专门的越狱社区和通用聊天机器人。采用概率度量、词汇多样性、压缩比及认知负荷等多种复杂度指标,发现越狱尝试的复杂度并不显著高于正常对话。这一现象在专业越狱社区与普通用户群体中均保持一致,表明攻击复杂度存在实际边界。时间序列分析显示,尽管用户攻击的毒性和复杂度长期稳定,但助手响应的毒性已下降,反映安全机制持续改进。复杂度分布未呈现幂律特性,进一步暗示越狱发展存在自然限制。研究挑战了攻击与防御不断升级的主流叙事,认为大模型安全演进受限于人类创造力,而防御措施仍在进步。结果强调学术越狱披露中的信息危害:若出现超越当前复杂度基准的高阶攻击,可能打破现有平衡,导致广泛危害而防御尚未适应。

原文摘要 · Abstract (English)

As large language models (LLMs) become increasingly deployed, understanding the complexity and evolution of jailbreaking strategies is critical for AI safety. We present a mass-scale empirical analysis of jailbreak complexity across over 2 million real-world conversations from diverse platforms, including dedicated jailbreaking communities and general-purpose chatbots. Using a range of complexity metrics spanning probabilistic measures, lexical diversity, compression ratios, and cognitive load indicators, we find that jailbreak attempts do not exhibit significantly higher complexity than normal conversations. This pattern holds consistently across specialized jailbreaking communities and general user populations, suggesting practical bounds on attack sophistication. Temporal analysis reveals that while user attack toxicity and complexity remains stable over time, assistant response toxicity has decreased, indicating improving safety mechanisms. The absence of power-law scaling in complexity distributions further points to natural limits on jailbreak development. Our findings challenge the prevailing narrative of an escalating arms race between attackers and defenders, instead suggesting that LLM safety evolution is bounded by human ingenuity constraints while defensive measures continue advancing. Our results highlight critical information hazards in academic jailbreak disclosure, as sophisticated attacks exceeding current complexity baselines could disrupt the observed equilibrium and enable widespread harm before defensive adaptation.

AI安全越狱攻击大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。