用Swin Transformer识别真假图像,跨数据集准确率达97%-99%。
Swin Transformer for Robust Differentiation of Real and Synthetic Images: Intra- and Inter-Dataset Analysis

- 基于Swin Transformer的分层结构捕捉图像局部与全局特征。
- 在三个数据集上测试,跨域识别准确率稳定在97%-99%。
- 适合需要高可靠性的数字图像真伪鉴别场景。
本研究旨在应对在RGB色彩空间中区分计算机生成图像(CGI)与真实数字图像日益严峻的挑战。针对现有分类方法在处理CGI复杂性和多样性方面的局限性,本文提出一种基于Swin Transformer的模型,用于精确区分自然图像与合成图像。该模型利用Swin Transformer的分层架构,捕获对区分CGI与自然图像至关重要的局部与全局特征。通过在三个独立数据集(CiFAKE、JSSSTU、Columbia)上的内部与跨数据集测试,评估了模型的鲁棒性与领域泛化能力。测试包括单个数据集(D1、D2、D3)及组合数据集(D1+D2+D3)。结果显示,该模型在所有数据集和测试场景中均保持97%-99%的高准确率,验证了其在检测CGI方面的有效性,展现出出色的鲁棒性与可靠性。研究结论表明,Swin Transformer模型在数字图像取证中具有重要潜力,尤其在区分合成与自然图像方面表现突出,且在多数据集上的优异表现证实其具备良好的领域泛化能力,是高精度、高可靠性图像分类任务中的有力工具。
原文摘要 · Abstract (English)
\textbf{Purpose} This study aims to address the growing challenge of distinguishing computer-generated imagery (CGI) from authentic digital images in the RGB color space. Given the limitations of existing classification methods in handling the complexity and variability of CGI, this research proposes a Swin Transformer-based model for accurate differentiation between natural and synthetic images. \textbf{Methods} The proposed model leverages the Swin Transformer's hierarchical architecture to capture local and global features crucial for distinguishing CGI from natural images. The model's performance was evaluated through intra-dataset and inter-dataset testing across three distinct datasets: CiFAKE, JSSSTU, and Columbia. The datasets were tested individually (D1, D2, D3) and in combination (D1+D2+D3) to assess the model's robustness and domain generalization capabilities. \textbf{Results} The Swin Transformer-based model demonstrated high accuracy, consistently achieving a range of 97-99\% across all datasets and testing scenarios. These results confirm the model's effectiveness in detecting CGI, showcasing its robustness and reliability in both intra-dataset and inter-dataset evaluations. \textbf{Conclusion} The findings of this study highlight the Swin Transformer model's potential as an advanced tool for digital image forensics, particularly in distinguishing CGI from natural images. The model's strong performance across multiple datasets indicates its capability for domain generalization, making it a valuable asset in scenarios requiring precise and reliable image classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。