arXiv:2506.10029cs.CRcs.AI2025-06

对比ChatGPT与Gemini的对抗攻击防御能力,揭示其安全短板。

Evaluation empirique de la sécurisation et de l'alignement de ChatGPT et Gemini: analyse comparative des vulnérabilités par expérimentations de jailbreaks

  • 通过实测多种越狱攻击,评估两模型安全防护机制。
  • 发现两者均易受特定提示注入攻击,存在显著漏洞。
  • 适合关注大模型安全风险的研究者与开发者参考。

大型语言模型(LLMs)正深刻改变数字应用,涵盖文本生成、图像创作、信息检索与代码开发等领域。2022年11月发布的ChatGPT迅速成为行业标杆,推动谷歌推出Gemini等竞品。然而,技术进步也带来新威胁:提示注入攻击、绕过监管机制(越狱)、虚假信息传播(幻觉)及深度伪造风险。本文对ChatGPT与Gemini的安全性与对齐程度进行比较分析,并通过实验构建越狱技术分类体系,揭示其在对抗攻击下的脆弱性。

原文摘要 · Abstract (English)

Large Language models (LLMs) are transforming digital usage, particularly in text generation, image creation, information retrieval and code development. ChatGPT, launched by OpenAI in November 2022, quickly became a reference, prompting the emergence of competitors such as Google's Gemini. However, these technological advances raise new cybersecurity challenges, including prompt injection attacks, the circumvention of regulatory measures (jailbreaking), the spread of misinformation (hallucinations) and risks associated with deep fakes. This paper presents a comparative analysis of the security and alignment levels of ChatGPT and Gemini, as well as a taxonomy of jailbreak techniques associated with experiments.

大模型安全越狱攻击LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。