arXiv:2601.22663cs.CVcs.AI2026-01

无需标注数据,通过对齐与解耦实现生成图像的自动概念溯源。

Unsupervised Synthetic Image Attribution: Alignment and Disentanglement

  • 基于对比自监督学习实现图像概念对齐,再用Infomax损失促进特征解耦。
  • 在AbC基准上,无监督方法性能超越有监督基线,意外取得更好效果。
  • 适合关注生成模型透明性与版权保护的研究者参考。

随着合成图像质量提升,识别模型生成图像背后的原始概念对于版权保护和模型透明性愈发重要。现有方法依赖大量带标注的合成图像与其训练源配对数据,但这类标注成本高昂。为此,本文探索无监督合成图像溯源方法,提出简单有效的「对齐与解耦」(Alignment and Disentanglement)框架:首先利用对比自监督学习(如MoCo、DINO)进行基础概念对齐;再通过Infomax损失增强表示解耦能力。该方法基于一个观察:对比自监督模型天然具备跨域对齐能力。我们将其形式化为关于交叉协方差的理论假设,并从典型相关分析目标分解角度解释了对齐与解耦如何逼近概念匹配过程。在真实世界基准AbC上,我们的无监督方法意外优于已有监督方法,为该难题提供了新视角。

原文摘要 · Abstract (English)

As the quality of synthetic images improves, identifying the underlying concepts of model-generated images is becoming increasingly crucial for copyright protection and ensuring model transparency. Existing methods achieve this attribution goal by training models using annotated pairs of synthetic images and their original training sources. However, obtaining such paired supervision is challenging, as it requires either well-designed synthetic concepts or precise annotations from millions of training sources. To eliminate the need for costly paired annotations, in this paper, we explore the possibility of unsupervised synthetic image attribution. We propose a simple yet effective unsupervised method called Alignment and Disentanglement. Specifically, we begin by performing basic concept alignment using contrastive self-supervised learning. Next, we enhance the model's attribution ability by promoting representation disentanglement with the Infomax loss. This approach is motivated by an interesting observation: contrastive self-supervised models, such as MoCo and DINO, inherently exhibit the ability to perform simple cross-domain alignment. By formulating this observation as a theoretical assumption on cross-covariance, we provide a theoretical explanation of how alignment and disentanglement can approximate the concept-matching process through a decomposition of the canonical correlation analysis objective. On the real-world benchmarks, AbC, we show that our unsupervised method surprisingly outperforms the supervised methods. As a starting point, we expect our intuitive insights and experimental findings to provide a fresh perspective on this challenging task.

图像溯源无监督学习生成模型特征解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。