提出PromptLA方法,高效验证文生图模型是否被篡改
PromptLA: Towards Integrity Verification of Black-box Text-to-Image Diffusion Models
- 通过特征分布的KL散度检测模型输出异常
- 在4个主流模型上平均AUC超0.96,比基线高0.2以上
- 低成本、抗图像后处理,适用于版权维权场景
尽管文本到图像(T2I)扩散模型生成质量优异,但其黑盒部署带来了重大监管挑战:恶意用户可通过微调模型生成违法内容,绕过现有安全机制。因此,验证T2I扩散模型的完整性至关重要。为此,考虑到生成模型输出中的随机性及与模型交互的高成本,我们基于生成图像特征分布的KL散度来识别模型篡改行为。提出一种基于学习自动机的新型提示选择算法(PromptLA),实现高效准确的验证。在四个先进T2I模型(如SDXL、FLUX.1)上的评估表明,该方法在完整性检测中平均AUC超过0.96,较基线提升超过0.2,展现出强有效性和泛化能力。此外,本方法成本更低且对图像级后处理具有鲁棒性。据我们所知,这是首个针对T2I扩散模型完整性验证的工作,为实际AI版权诉讼建立了可量化的标准。
原文摘要 · Abstract (English)
Despite the impressive synthesis quality of text-to-image (T2I) diffusion models, their black-box deployment poses significant regulatory challenges: Malicious actors can fine-tune these models to generate illegal content, circumventing existing safeguards through parameter manipulation. Therefore, it is essential to verify the integrity of T2I diffusion models. To this end, considering the randomness within the outputs of generative models and the high costs in interacting with them, we discern model tampering via the KL divergence between the distributions of the features of generated images. We propose a novel prompt selection algorithm based on learning automaton (PromptLA) for efficient and accurate verification. Evaluations on four advanced T2I models (e.g., SDXL, FLUX.1) demonstrate that our method achieves a mean AUC of over 0.96 in integrity detection, exceeding baselines by more than 0.2, showcasing strong effectiveness and generalization. Additionally, our approach achieves lower cost and is robust against image-level post-processing. To the best of our knowledge, this paper is the first work addressing the integrity verification of T2I diffusion models, which establishes quantifiable standards for AI copyright litigation in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。