arXiv:2506.04036cs.CRcs.AI2025-06被引 1

98.8%的自定义GPT存在指令泄露风险,威胁隐私与安全。

Privacy and Security Threat for OpenAI GPTs

  • 设计三阶段攻击方法,测试真实GPT的指令泄露漏洞。
  • 超98%自定义GPT易受攻击,77.5%有防御仍不安全。
  • 发现738个GPT窃取用户对话,8个存在无关数据访问行为。

大型语言模型(LLMs)具备强大的信息处理能力,广泛应用于聊天机器人。OpenAI提供平台供开发者创建自定义GPT,扩展ChatGPT功能并集成外部服务。自2023年11月发布以来,已创建超过300万个自定义GPT。然而,这一庞大生态系统也隐藏着安全与隐私风险。对开发者而言,对抗性提示可导致指令泄露,侵犯知识产权;对用户而言,自定义GPT或第三方服务的不当数据访问行为引发严重隐私担忧。为系统评估真实场景下威胁范围,我们设计三阶段指令泄露攻击,针对不同防御等级的GPT进行测试。在10,000个真实自定义GPT上的广泛实验表明,超过98.8%的GPT可通过一种或多种对抗性提示被攻破,剩余一半还可通过多轮对话攻击。我们还构建框架评估防御策略有效性,并识别异常行为。结果表明,77.5%具有防御机制的GPT仍易受基础攻击。此外,发现738个GPT收集用户对话信息,其中8个表现出与其功能无关的数据访问行为。研究警示开发者应强化指令防御,提醒用户关注基于LLM应用的数据隐私风险。

原文摘要 · Abstract (English)

Large language models (LLMs) demonstrate powerful information handling capabilities and are widely integrated into chatbot applications. OpenAI provides a platform for developers to construct custom GPTs, extending ChatGPT's functions and integrating external services. Since its release in November 2023, over 3 million custom GPTs have been created. However, such a vast ecosystem also conceals security and privacy threats. For developers, instruction leaking attacks threaten the intellectual property of instructions in custom GPTs through carefully crafted adversarial prompts. For users, unwanted data access behavior by custom GPTs or integrated third-party services raises significant privacy concerns. To systematically evaluate the scope of threats in real-world LLM applications, we develop three phases instruction leaking attacks target GPTs with different defense level. Our widespread experiments on 10,000 real-world custom GPTs reveal that over 98.8% of GPTs are vulnerable to instruction leaking attacks via one or more adversarial prompts, and half of the remaining GPTs can also be attacked through multiround conversations. We also developed a framework to assess the effectiveness of defensive strategies and identify unwanted behaviors in custom GPTs. Our findings show that 77.5% of custom GPTs with defense strategies are vulnerable to basic instruction leaking attacks. Additionally, we reveal that 738 custom GPTs collect user conversational information, and identified 8 GPTs exhibiting data access behaviors that are unnecessary for their intended functionalities. Our findings raise awareness among GPT developers about the importance of integrating specific defensive strategies in their instructions and highlight users' concerns about data privacy when using LLM-based applications.

大模型安全隐私泄露GPT防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。