arXiv:2607.12165cs.CV2026-07

对比合成图像与真实图像差异,提出安全使用合成数据的评估方法。

Data Safety: Synthetic Data Quality Analysis Using CIFAKE Dataset

论文配图:Data Safety: Synthetic Data Quality Analysis Using CIFAKE Dataset
图 1 · 摘自论文原文
  • 从高维特征、颜色统计和训练过程三方面分析合成与真实图像差异
  • 验证不同混合场景下合成数据对模型性能的影响
  • 为未知质量合成数据提供安全评估与应用策略,适合模型开发者

近年来,高性能图像分类模型的社会化应用迅速扩展。尽管这些模型需要大量训练数据以提升性能,但获取充足的真实图像往往不切实际。为弥补这一不足,合成数据的使用日益广泛。然而,合成图像并不一定等同于真实图像用于训练。本研究从高维特征空间、颜色空间的低层统计特性以及模型训练过程三个角度,系统分析了不同生成方法产生的两类合成图像与真实图像之间的差异。同时,通过实验验证在现实数据混合场景中合成数据的使用方式,从而提出一种针对未知质量合成图像的初步评估与安全融入训练的策略。该研究旨在提升利用合成图像的图像分类模型的可靠性与安全性。

原文摘要 · Abstract (English)

Recently, the societal implementation of high-performance image classification models has expanded rapidly. While these models require vast amounts of training data to improve performance, securing sufficient real images is often impractical. As a means to compensate for this shortage, the use of synthetic data is becoming widespread. However, synthetic images are not necessarily equivalent to real images for training purposes. This study systematically analyzes the differences between two types of synthetic images created by different generation methods and real images from three perspectives: high-dimensional feature space, low-level statistics in color space, and the model training process. Furthermore, it experimentally verifies how synthetic data should be utilized by considering realistic data mixing scenarios. This enables the proposal of an evaluation and application strategy for performing preliminary assessments on synthetic images of unknown quality and safely incorporating them into training. This research aims to contribute to enhancing the reliability and safety of image classification models utilizing synthetic images.

合成数据图像质量模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。