arXiv:2607.13685cs.CV2026-07

无需训练即可精准追踪生成图像的源模型,区分高度相似的变体。

DNA: Dual-stage Native Attribution for Generated Image Source Tracing

论文配图:DNA: Dual-stage Native Attribution for Generated Image Source Tracing
图 1 · 摘自论文原文
  • 分两阶段利用生成模型层级特征:先粗筛家族,再精判具体模型。
  • 在30,000张图像上实现89.11%准确率,远超随机猜测(<1%)。
  • 适用于未知源和难区分的同类变体,适合数字取证场景。

图像生成技术快速发展,催生了大量同家族变体,溯源任务对数字取证愈发重要。现有主动方法依赖水印或模型修改,可能影响画质且部署受限;被动方法常依赖大规模监督训练或单一重建信号,难以应对未知源及高度相似变体。我们发现生成模型隐空间中的归属信号具有层级特性:VAE层反映家族共性,主干层捕捉变体特异性。基于此,提出双阶段原生溯源框架DNA,无需额外训练。粗粒度阶段采用自编码器双重重建(AEDR)实现高效开集家族级筛选;细粒度阶段通过原生预测一致性(NPC)比较同一变体在多噪声水平下语义条件下的原生预测误差,以归一化校准得分判定来源。为系统评估,构建DNA-30K基准数据集,包含24个候选模型、6个家族、30,000张生成图像(覆盖去噪扩散与流匹配),以及非候选生成与自然图像作为未知源。实验表明,DNA在端到端任务中达89.11%准确率,远超随机猜测(<1%),且相比最强基线提升33.81%,即使使用AEDR作粗筛也表现优异。

原文摘要 · Abstract (English)

The rapid evolution of image generation has produced numerous within-family variants, making source-model attribution of suspect images increasingly important for digital forensics. Existing proactive methods rely on watermark embedding or model modification, which may degrade visual quality and limit deployment flexibility. Passive methods often rely on large-scale supervised training or a single reconstruction signal, limiting their ability to handle unknown sources and distinguish highly similar within-family variants. We observe that attribution signals in latent generative models are naturally stratified across architectural levels: VAE-level cues reflect family-shared information, whereas backbone-level cues capture variant-specific behaviors. Motivated by this insight, we propose Dual-stage Native Attribution (DNA), a coarse-to-fine framework that follows this hierarchy without additional neural-network training. The coarse-grained stage uses Autoencoder Double-Reconstruction (AEDR) for efficient open-set family-level screening. The fine-grained stage performs closed-set model-level attribution with Native Prediction Consistency (NPC), which compares native prediction errors of within-family variants across multiple noise levels under semantic conditioning and attributes the source via normalized calibrated scores. To enable systematic evaluation, we construct DNA-30K, a benchmark for within-family variant attribution under open-set family-level evaluation. It comprises 30,000 images generated by 24 candidate models across six families spanning both denoising diffusion and flow matching, plus non-candidate generated and natural images as unknown sources. Experiments show that DNA achieves 89.11% end-to-end attribution accuracy on a task where random guessing accuracy is below 1% and outperforms the strongest baseline by 33.81% even when AEDR is used as the coarse-grained stage.

图像溯源生成模型数字取证无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。