arXiv:2506.22802cs.LGcs.CR2025-06ICCV被引 7

用黎曼几何分析生成模型的指纹,更好区分不同模型和合成数据。

Riemannian-Geometric Fingerprints of Generative Models

  • 基于黎曼几何定义生成模型的指纹,用测地距离替代欧氏距离。
  • 在4个数据集、27种模型架构上显著提升模型溯源准确率。
  • 适用于视觉与图文模型,对未见数据集有强泛化能力。

生成模型(GMs)的快速普及催生了模型归属与指纹识别的需求。服务提供商需保护知识产权,用户与执法机构则需验证内容来源以确保可信度。同时,模型生成数据回流训练引发“模型坍塌”风险,亟需区分真实与合成数据。然而,现有研究缺乏统一的理论框架来定义和分析生成模型的指纹。为此,本文提出一种基于黎曼几何的新方法:从数据中学习黎曼度量,用测地距离和基于kNN的黎曼质心取代欧氏距离与最近邻搜索,从而在非欧空间中定义模型指纹。该方法应用于新型梯度算法,可在64×64与256×256分辨率下,覆盖4个数据集、27种模型架构及视觉与视觉-语言双模态场景中,有效区分多种生成模型。实验表明,新方法显著提升模型归属性能,并具备对未见数据集、模型类型与模态的良好泛化能力,验证了其实际有效性。

原文摘要 · Abstract (English)

Recent breakthroughs and rapid integration of generative models (GMs) have sparked interest in the problem of model attribution and their fingerprints. For instance, service providers need reliable methods of authenticating their models to protect their IP, while users and law enforcement seek to verify the source of generated content for accountability and trust. In addition, a growing threat of model collapse is arising, as more model-generated data are being fed back into sources (e.g., YouTube) that are often harvested for training ("regurgitative training"), heightening the need to differentiate synthetic from human data. Yet, a gap still exists in understanding generative models' fingerprints, we believe, stemming from the lack of a formal framework that can define, represent, and analyze the fingerprints in a principled way. To address this gap, we take a geometric approach and propose a new definition of artifact and fingerprint of GMs using Riemannian geometry, which allows us to leverage the rich theory of differential geometry. Our new definition generalizes previous work (Song et al., 2024) to non-Euclidean manifolds by learning Riemannian metrics from data and replacing the Euclidean distances and nearest-neighbor search with geodesic distances and kNN-based Riemannian center of mass. We apply our theory to a new gradient-based algorithm for computing the fingerprints in practice. Results show that it is more effective in distinguishing a large array of GMs, spanning across 4 different datasets in 2 different resolutions (64 by 64, 256 by 256), 27 model architectures, and 2 modalities (Vision, Vision-Language). Using our proposed definition significantly improves the performance on model attribution, as well as a generalization to unseen datasets, model types, and modalities, suggesting its practical efficacy.

生成模型指纹识别黎曼几何模型溯源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。