arXiv:2607.22808cs.CVcs.AI2026-07中稿 · the DLMMDD Challen…

用语义与数学取证融合方法,精准识别生成图像来源。

Hybrid Semantic and Spectral Ensemble for Robust Synthetic Image Source Attribution

  • 双分支集成:语义学习+数学特征提取,提升鲁棒性。
  • 在55%退化图像上达到95.60%准确率,效果稳定。
  • 纯CPU运行,适合真实场景部署,效率高。

文本到图像(T2I)模型的快速发展催生了对合成图像来源归属(SIA)的强需求。当前核心挑战在于训练图像与实际部署图像间的分布偏移,后者常经未知后处理如JPEG压缩、模糊等。本文针对ICANN 2026 DLMMDD挑战赛提出一种双分支集成框架,融合语义深度学习与数学取证特征提取。语义分支采用加权指数移动平均(EMA)与标签平滑正则化的EfficientNet-B0;取证分支从高通噪声残差中提取126个数学特征(含SVD谱轮廓和局部二值模式),经截断SVD压缩后由XGBoost分类。在包含10个生成器的数据集上测试,其中55%为退化图像,本方法取得95.60%私有排行榜准确率。整个流程计算高效,无需GPU加速,标准CPU端到端完成仅需不足6.5小时,凸显数学取证在真实场景中的实用性与可扩展性。

原文摘要 · Abstract (English)

The rapid advancement of text-to-image (T2I) models has necessitated robust Synthetic Image Source Attribution (SIA) methodologies. A critical challenge in SIA is the distribution shift between pristine training images and real-world deployed images, which undergo unknown post-processing operations such as JPEG compression and blurring. In this work, proposed for the DLMMDD Challenge at ICANN 2026, we introduce a dual-branch ensemble framework fusing Semantic Deep Learning with Mathematical Forensic Feature Extraction. The semantic branch employs EfficientNet-B0 regularized with Exponential Moving Averaging (EMA) and Label Smoothing. The forensic branch extracts 126 mathematical features -- including SVD spectral profiles and Local Binary Patterns -- from high-pass noise residuals, compressed via Truncated SVD and classified with XGBoost. Evaluated on a dataset of 10 generators where 55% of the test set is degraded, our approach achieves a private leaderboard accuracy of 95.60%. Furthermore, the entire pipeline is highly computationally efficient, requiring no GPU acceleration and executing end-to-end on a standard CPU in under 6.5 hours, highlighting the practicality and scalability of mathematical forensics for real-world deployment.

图像溯源数学取证生成模型高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。