用Swin Transformer识别合成图像,跨数据集表现更稳定。
Swin Transformer for Robust CGI Images Detection: Intra- and Inter-Dataset Analysis across Multiple Color Spaces
- 基于Swin Transformer的层级结构捕捉局部与全局特征。
- 在三个数据集上,RGB色彩空间准确率最高,达98.7%。
- 适合数字图像取证、AI生成内容检测等场景使用。
本研究针对计算机生成图像(CGI)与真实数字图像区分难题,探索了RGB、YCbCr和HSV三种色彩空间下的检测方法。针对现有分类方法在处理CGI复杂性与变异性方面的局限,提出基于Swin Transformer的模型,利用其分层架构捕获局部与全局特征以区分自然与合成图像。在CiFAKE、JSSSTU和Columbia三个数据集上进行跨数据集与跨色彩空间测试,评估模型在单一数据集(D1、D2、D3)及合并数据集(D1+D2+D3)上的鲁棒性与域泛化能力。通过数据增强缓解数据不平衡问题,并使用t-SNE可视化验证特征可分性。实验表明,所有色彩空间中,RGB方案性能最优,单个数据集最高准确率达98.7%。进一步对比VGG-19与ResNet-50等CNN模型,结果证实该模型在跨数据集检测中具备更强鲁棒性与可靠性,展现出在数字图像取证领域的应用潜力。
原文摘要 · Abstract (English)
This study aims to address the growing challenge of distinguishing computer-generated imagery (CGI) from authentic digital images across three different color spaces; RGB, YCbCr, and HSV. Given the limitations of existing classification methods in handling the complexity and variability of CGI, this research proposes a Swin Transformer based model for accurate differentiation between natural and synthetic images. The proposed model leverages the Swin Transformer's hierarchical architecture to capture local and global features for distinguishing CGI from natural images. Its performance was assessed through intra- and inter-dataset testing across three datasets: CiFAKE, JSSSTU, and Columbia. The model was evaluated individually on each dataset (D1, D2, D3) and on the combined datasets (D1+D2+D3) to test its robustness and domain generalization. To address dataset imbalance, data augmentation techniques were applied. Additionally, t-SNE visualization was used to demonstrate the feature separability achieved by the Swin Transformer across the selected color spaces. The model's performance was tested across all color schemes, with the RGB color scheme yielding the highest accuracy for each dataset. As a result, RGB was selected for domain generalization analysis and compared with other CNN-based models, VGG-19 and ResNet-50. The comparative results demonstrate the proposed model's effectiveness in detecting CGI, highlighting its robustness and reliability in both intra-dataset and inter-dataset evaluations. The findings of this study highlight the Swin Transformer model's potential as an advanced tool for digital image forensics, particularly in distinguishing CGI from natural images. The model's strong performance indicates its capability for domain generalization, making it a valuable asset in scenarios requiring precise and reliable image classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。