arXiv:2608.13671cs.CV2026-08

无需训练即可还原图像生成提示,且每条恢复内容都有图像证据支持。

PROVE: Training-Free Prompt Recovery using Verifiable Evidence

论文配图:PROVE: Training-Free Prompt Recovery using Verifiable Evidence
图 1 · 摘自论文原文
  • 通过可验证的场景描述组合还原提示,而非优化词元序列。
  • 在多个数据集上超越现有方法,在图像相似性和图文对齐上表现更优。
  • 适合关注版权保护与生成内容安全的研究者和创作者。

现代文本到图像模型能从自然语言提示生成高度逼真的图像,而提示反演技术的发展使得从生成结果中恢复提示变得日益可行,引发版权保护与内容所有权的新担忧。随着提示市场兴起,恢复出的提示可能被用于未经授权复制和分发受版权保护的作品,以及暴露艺术家在生成内容中编码的创作配方。现有提示反演方法依赖梯度优化、自回归描述或强化学习,但优化方法常产生不可读提示,描述方法会虚构未经验证细节,强化学习方法则易过拟合特定生成器并引入评估循环问题。我们提出PROVE(Prompt Recovery with Verified Evidence),一种无需训练、黑盒式的提示反演攻击,通过组合可验证的场景描述来重构提示,适用于原始受版权保护作品及AI生成内容。恢复的提示完全可审计,每一项恢复内容均有明确图像证据支撑,并通过精度约束下的召回最大化目标进行形式化。在MS-COCO、Flickr30K和Lexica数据集上,使用最先进的文本到图像生成器,PROVE在图像相似性(DINO、LPIPS)和图文对齐(CLIP)方面持续优于优化、描述和强化学习基线方法,且无需训练、生成器访问或微调,展现出更强且更实用的提示反演能力。

原文摘要 · Abstract (English)

Modern text-to-image models can generate highly realistic images from natural-language prompts, while recent advances in prompt inversion have made it increasingly feasible to recover those prompts from generated outputs, raising new concerns for copyright protection and content ownership. As prompt marketplaces emerge, recovered prompts can enable both the unauthorized reproduction and redistribution of copyrighted creative works, and the exposure of the prompts that encode an artist's creative recipe in AI-generated content. Existing prompt inversion methods rely on gradient-based optimization, autoregressive captioning, or reinforcement learning. However, optimization-based methods often produce unreadable prompts, captioning methods hallucinate unverified details, and RL-based approaches frequently overfit to specific generators while introducing evaluation circularity. We introduce PROVE (Prompt Recovery with Verified Evidence), a training-free, black-box prompt inversion attack that reconstructs prompts by composing verifiable scene descriptions rather than optimizing token sequences, targeting both original copyrighted works and AI-generated content. The resulting prompts are fully auditable, with every recovered claim grounded in explicit image evidence, and are formalized through a precision-constrained recall maximization objective. Across MS-COCO, Flickr30K, and Lexica, using state-of-the-art text-to-image generators, PROVE consistently outperforms optimization, captioning, and RL-based baselines on image similarity (DINO, LPIPS) and text-image alignment (CLIP), without any training, generator access, or fine-tuning, demonstrating a stronger and more practical prompt inversion attack.

提示反演版权保护可验证性生成安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。