arXiv:2505.08148cs.CRcs.AI2025-05被引 7

分析1.5万款自定义GPT,发现95%缺乏基本安全防护。

A Large-Scale Empirical Analysis of Custom GPTs' Vulnerabilities in the OpenAI Ecosystem

  • 从1.5万款自定义GPT中实证检测七类安全漏洞
  • 超96%存在角色扮演攻击漏洞,92%有系统提示泄露风险
  • 揭示基础模型缺陷被定制模型放大,适合关注AI安全的开发者

数百万用户使用由主流模型提供商开发的生成式预训练变换器(GPT)模型完成各类任务。为支持更丰富的交互与定制化,如OpenAI等平台现允许开发者通过专用仓库或应用商店创建并发布定制化模型实例,称为自定义GPT,使用户可浏览并交互于满足特定需求的专用应用。然而,随着自定义GPT普及,其安全漏洞问题日益突出。现有研究多为理论探讨,缺乏大规模、统计严谨的实证评估。本研究分析了14,904款自定义GPT,评估其在角色扮演攻击、系统提示泄露、钓鱼内容生成、恶意代码合成等七类可利用威胁下的脆弱性,覆盖不同类别与流行度层级。引入多指标排名体系,探究流行度与安全风险的关系。结果表明,超过95%的自定义GPT缺乏充分安全防护。最常见漏洞包括角色扮演漏洞(96.51%)、系统提示泄露(92.20%)和钓鱼内容生成(91.22%)。此外,我们证明OpenAI基础模型存在固有安全弱点,常被自定义GPT继承或放大。研究凸显亟需加强安全措施与内容审核,以保障GPT应用的安全部署。

原文摘要 · Abstract (English)

Millions of users leverage generative pretrained transformer (GPT)-based language models developed by leading model providers for a wide range of tasks. To support enhanced user interaction and customization, many platforms-such as OpenAI-now enable developers to create and publish tailored model instances, known as custom GPTs, via dedicated repositories or application stores. These custom GPTs empower users to browse and interact with specialized applications designed to meet specific needs. However, as custom GPTs see growing adoption, concerns regarding their security vulnerabilities have intensified. Existing research on these vulnerabilities remains largely theoretical, often lacking empirical, large-scale, and statistically rigorous assessments of associated risks. In this study, we analyze 14,904 custom GPTs to assess their susceptibility to seven exploitable threats, such as roleplay-based attacks, system prompt leakage, phishing content generation, and malicious code synthesis, across various categories and popularity tiers within the OpenAI marketplace. We introduce a multi-metric ranking system to examine the relationship between a custom GPT's popularity and its associated security risks. Our findings reveal that over 95% of custom GPTs lack adequate security protections. The most prevalent vulnerabilities include roleplay-based vulnerabilities (96.51%), system prompt leakage (92.20%), and phishing (91.22%). Furthermore, we demonstrate that OpenAI's foundational models exhibit inherent security weaknesses, which are often inherited or amplified in custom GPTs. These results highlight the urgent need for enhanced security measures and stricter content moderation to ensure the safe deployment of GPT-based applications.

AI安全GPT漏洞大模型风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。