arXiv:2508.05691cs.CRcs.AI2025-08被引 3

通过隐藏验证任务提升生成图像溯源的抗攻击能力

SPRINT: Robust Model Attribution of Generated Images via Secret Pixel Reconstruction

  • 用秘密像素重建机制隐藏指纹验证方式,使攻击者无法定位目标
  • 在FFHQ数据集上对12个模型达99.17%准确率,对抗攻击成功率低于1%
  • 适合需要高安全性的图像溯源场景,如内容审核与版权追踪

检测AI生成图像的来源模型是日益重要的责任问题。现有指纹技术通过识别图像中不可察觉的、每个模型独有的模式实现高精度检测,但在面对自适应攻击时极脆弱——攻击者可利用技术知识干扰指纹以逃避检测。本文提出SPRINT(秘密像素重建指纹),一种专为抵御自适应攻击设计的新方法。不同于传统依赖公开可见模式的指纹,SPRINT基于秘密定义隐藏的重建目标,使验证任务本身保持私有。因此,攻击者在验证阶段无法知晓需破解的任务,从而难以发动针对性攻击。实验表明,SPRINT在理想条件下仍保持高精度:在FFHQ数据集上,对12个不同模型组合的闭世界测试中达到99.17%准确率,在6个同架构相近检查点的更难组合中达98.83%;同时将自适应移除和伪造攻击成功率降至1%以下。在开放世界设置下,对相同模型池仍维持99.30%的AUROC。结果证明,将验证任务私有化可显著提高对自适应攻击的鲁棒性,同时在干净环境下保持优异性能。

原文摘要 · Abstract (English)

Detecting the source model of AI-generated images is a growing accountability problem. AI fingerprinting techniques address this by detecting imperceptible patterns in the images that are unique to each model, achieving high detection accuracy under ideal conditions. However, recent research has shown that image fingerprints are extremely brittle to adaptive attacks, where knowledge of the technique can be exploited to perturb the fingerprints and evade detection. We present SPRINT (Secret Pixel Reconstruction fingerprinting), a novel model attribution method specifically designed to provide robustness to adaptive attacks. As opposed to existing fingerprinting, which focuses on publicly discoverable patterns in the image, SPRINT relies on a secret to define hidden reconstruction targets, thus keeping the verification task itself private. As a result, the attacker can no longer see the task that the verifier solves at verification time, protecting the information exploited by the attacks. Our results show that SPRINT achieves high closed-world accuracy while remaining robust to adaptive attacks: on the FFHQ dataset, SPRINT reaches 99.17% clean accuracy on a diverse 12-model pool and 98.83% on a harder pool of 6 close checkpoints of the same model architecture, while reducing adaptive removal and forgery attack success rates to 1% or below. When the same pool of close model checkpoints is considered an open world, SPRINT maintains high accuracy with an AUROC of 99.30%. These findings show that the approach of privatizing the verification task can make adaptive evasion substantially harder while maintaining performance in the clean setting.

图像溯源对抗鲁棒性隐私指纹生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。