用强化学习偷取文本生成图像的提示模板,效率超现有方法90%以上。
Reinforcement Learning-Based Prompt Template Stealing for Text-to-Image Models
- 将提示窃取建模为序列决策问题,用相似度反馈做奖励
- 仅需少量样本图即可还原模板,攻击成本降至13%以下
- 能跨风格泛化,适合关注生成模型安全的研究者
多模态大语言模型(MLLMs)已革新文生图流程,使设计师能以空前速度创造视觉概念。这一进展催生了活跃的提示交易市场,其中精心设计的提示可生成特定风格。尽管商业价值高,但提示本身存在严重安全隐患:可能被窃取。本文揭示该风险并提出RLStealer——一种基于强化学习的提示模板逆向框架,仅需少量示例图像即可恢复模板。RLStealer将窃取过程视为序列决策问题,采用多种基于相似度的反馈信号作为奖励函数,有效探索提示空间。在公开基准上的实验表明,其性能达当前最优,且总攻击成本低于现有基线的13%。进一步分析显示,该方法能有效跨风格泛化,高效窃取未见过的提示模板。本研究揭示了提示交易中的紧迫安全威胁,并为新兴的MLLMs市场建立防护标准奠定基础。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have transformed text-to-image workflows, allowing designers to create novel visual concepts with unprecedented speed. This progress has given rise to a thriving prompt trading market, where curated prompts that induce trademark styles are bought and sold. Although commercially attractive, prompt trading also introduces a largely unexamined security risk: the prompts themselves can be stolen. In this paper, we expose this vulnerability and present RLStealer, a reinforcement learning based prompt inversion framework that recovers its template from only a small set of example images. RLStealer treats template stealing as a sequential decision making problem and employs multiple similarity based feedback signals as reward functions to effectively explore the prompt space. Comprehensive experiments on publicly available benchmarks demonstrate that RLStealer gets state-of-the-art performance while reducing the total attack cost to under 13% of that required by existing baselines. Our further analysis confirms that RLStealer can effectively generalize across different image styles to efficiently steal unseen prompt templates. Our study highlights an urgent security threat inherent in prompt trading and lays the groundwork for developing protective standards in the emerging MLLMs marketplace.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。