arXiv:2502.09974cs.AIcs.CR2025-02被引 8

检测大模型是否使用了特定系统提示,保护提示隐私。

Has My System Prompt Been Used? Large Language Model Prompt Membership Inference

  • 通过对比输出分布差异,判断模型是否使用目标提示。
  • 微小提示变化也会导致响应分布显著不同。
  • 适合关注提示安全与模型隐私的研究者。

提示工程已成为优化大语言模型(LLMs)以适应特定应用的强大技术,加速原型设计并提升性能,也引发了社区对专有系统提示保护的关注。本文从成员推断的角度探索提示隐私的新视角,提出 Prompt Detective——一种基于统计检验的方法,可可靠判断第三方语言模型是否使用了给定的系统提示。该方法通过比较不同系统提示对应模型输出的分布差异实现推断。在多种语言模型上的大量实验表明,即使系统提示的细微改动也会在响应分布中体现显著差异,从而实现具有统计显著性的提示使用验证。本工作揭示了提示隐私面临的新风险。

原文摘要 · Abstract (English)

Prompt engineering has emerged as a powerful technique for optimizing large language models (LLMs) for specific applications, enabling faster prototyping and improved performance, and giving rise to the interest of the community in protecting proprietary system prompts. In this work, we explore a novel perspective on prompt privacy through the lens of membership inference. We develop Prompt Detective, a statistical method to reliably determine whether a given system prompt was used by a third-party language model. Our approach relies on a statistical test comparing the distributions of two groups of model outputs corresponding to different system prompts. Through extensive experiments with a variety of language models, we demonstrate the effectiveness of Prompt Detective for prompt membership inference. Our work reveals that even minor changes in system prompts manifest in distinct response distributions, enabling us to verify prompt usage with statistical significance.

提示隐私成员推断LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。