为生成图像防窃取,添加不可见扰动保护输出质量
Safeguarding Text-to-Image Generative Models Against Unauthorized Knowledge Distillation

- 用频率感知生成器在单次遍历中注入不可见扰动
- 在保持视觉质量的同时显著降低盗用模型的训练效果
- 支持用户指定扰动强度,适合大规模图像发布场景
闭源生成服务通过查询式API提供输出,但攻击者可反复调用并收集大量合成图像,用于训练替代模型实现能力复制。为应对这一威胁,防御方案需兼顾图像保真度、扰动幅度可控性与大规模部署效率。本文提出WaveGuard,一种基于生成器的单次处理防护框架,在用户指定的扰动预算下,通过频率感知的扰动生成器注入结构化、不可察觉的噪声。该方法在良性用户眼中保持感知可用性,同时大幅降低受保护图像对未经授权的学生模型的训练价值。在与WikiArt相关的合成输出窃取场景下,WaveGuard实现了出色的效用-保真-效率平衡,具备明确的不可察觉性控制和显著的防护效率提升。
原文摘要 · Abstract (English)
Closed-weight generative services are increasingly deployed through query-based APIs, where users can obtain generated outputs while model parameters remain inaccessible. However, such deployment does not prevent model stealing: an attacker can repeatedly query the service, collect large volumes of released synthetic images, and use them as training data for a private substitute model. This query-output-driven process enables unauthorized knowledge distillation and capability replication without direct access to the original weights. To mitigate this threat, a practical defense should preserve the visual fidelity of released images, provide explicit control over perturbation magnitude, and scale efficiently to large-volume output release. We present WaveGuard, a single-pass, generator-based protection framework that safeguards released synthetic images under a user-specified perturbation budget. WaveGuard employs a frequency-aware perturbation generator to inject structured, imperceptible perturbations that maintain perceptual utility for benign viewers while reducing the usefulness of protected images as training data for unauthorized student models. Extensive experiments under WikiArt-related synthetic-output distillation settings show that WaveGuard achieves a favorable efficacy--fidelity--efficiency trade-off, with explicit imperceptibility control and substantial gains in protection efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。