解决伪造人脸生成器未知来源识别难题,实现精准溯源与持续发现。
Face-Trace: Open-Set Attribution and Progressive Discovery of Synthetic Face Generators

- 用冻结的I-JEPA特征+能量判别,区分已知与未知生成器。
- 未知样本聚类效果佳:调整兰德指数0.81,纯度达87.74%。
- 支持增量学习,可随时间扩展新生成器识别能力。
生成式AI使伪造人脸日益逼真,给多媒体取证带来挑战。传统方法多在封闭集下工作,假设测试样本仅来自训练中见过的生成器,这与现实不符。本文提出Face-Trace,一个开放集伪造人脸溯源框架,融合已知生成器分类、基于能量的拒绝机制和未知生成器发现。使用冻结的I-JEPA嵌入进行分类,拒绝样本则通过投影I-JEPA特征与互补取证特征组合,聚类形成有意义的未知生成器组。在增量场景下,该方法可逐步扩展识别空间。WILD数据集实验显示,封闭集准确率达96.73%,拒绝平衡准确率71.25%;聚类指标包括调整兰德指数0.81、归一化互信息0.90、整体纯度87.74%。增量设置下最终纯度达99.23%,跨数据集实验表明其具备泛化能力。
原文摘要 · Abstract (English)
Recent advances in generative Artificial Intelligence have made synthetic face images increasingly realistic, creating new challenges for multimedia forensics. Source attribution methods should identify the generator of an image when the source is known, but also handle samples produced by unseen models. Most existing approaches, however, address synthetic face attribution in a closed-set setting, assuming that test samples can only originate from generators observed during training. This assumption does not hold in real-world scenarios, where new generators continuously appear and detecting an image as unknown is not sufficient, since rejected samples should also be organized according to their underlying sources. We introduce Face-Trace, a pipeline for open-set synthetic face source attribution that combines known generator classification, energy-based rejection, and unknown generator discovery. A classifier trained on frozen I-JEPA embeddings attributes known generators, while rejected samples are represented by combining projected I-JEPA features with complementary forensic traces and grouped to identify coherent sets of samples produced by unknown generators. We also extend the discovery stage to an incremental scenario, where rejected samples arrive over time. Experiments on the WILD dataset show 96.73% closed-set attribution accuracy, while rejection reaches 71.25% balanced accuracy and rejected samples are clustered into meaningful unknown-generator groups, with an Adjusted Rand Index of 0.81, a Normalized Mutual Information of 0.90, and an overall purity of 87.74%. In the incremental setting, the discovered generator space is progressively extended while maintaining a final purity of 99.23%, and cross-dataset experiments suggest that the pipeline can operate beyond the original data distribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。