arXiv:2508.16347cs.CRcs.AI2025-08EMNLP被引 11

揭露大模型越狱测试的虚假安全,指出其真实犯罪知识缺失

Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs

  • 用问答任务分离越狱技术,检验模型是否真掌握危险知识
  • 发现越狱成功率与真实危害知识掌握度严重不匹配
  • 适合关注大模型真实滥用风险的研究者和安全评估者

随着大语言模型(LLMs)的发展,大量研究揭示了其对越狱攻击的脆弱性。尽管这些工作推动了大模型的安全对齐进展,但尚不清楚模型是否真正内化了应对现实犯罪的常识,还是仅能模仿有毒语言模式。这种模糊性引发担忧:越狱成功率可能源于被越狱大模型与评判模型之间的幻觉循环。通过解耦越狱技术使用,我们构建了知识密集型问答任务,从危险知识掌握、有害任务规划能力和有害性判断鲁棒性三方面考察大模型的真实滥用威胁。实验显示,越狱成功率与有害知识掌握程度存在明显脱节,现有以大模型为评判者的框架往往将有害性判断锚定在有毒语言模式上。本研究揭示了当前大模型安全评估与真实世界威胁潜力之间的差距。

原文摘要 · Abstract (English)

With the development of Large Language Models (LLMs), numerous efforts have revealed their vulnerabilities to jailbreak attacks. Although these studies have driven the progress in LLMs' safety alignment, it remains unclear whether LLMs have internalized authentic knowledge to deal with real-world crimes, or are merely forced to simulate toxic language patterns. This ambiguity raises concerns that jailbreak success is often attributable to a hallucination loop between jailbroken LLM and judger LLM. By decoupling the use of jailbreak techniques, we construct knowledge-intensive Q\&A to investigate the misuse threats of LLMs in terms of dangerous knowledge possession, harmful task planning utility, and harmfulness judgment robustness. Experiments reveal a mismatch between jailbreak success rates and harmful knowledge possession in LLMs, and existing LLM-as-a-judge frameworks tend to anchor harmfulness judgments on toxic language patterns. Our study reveals a gap between existing LLM safety assessments and real-world threat potential.

大模型安全越狱攻击真实威胁

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。