发现生成图像检测的‘最不利区间’:中等复杂度数据集让检测器最易胜出。
The Unwinnable Arms Race of AI Image Detection
- 用柯尔莫哥洛夫复杂度衡量数据结构,揭示复杂度对检测的影响
- 极简单或极复杂的数据集都难被检测,中等复杂度最易暴露生成痕迹
- 适合关注生成内容安全、对抗检测机制的研究者阅读
生成式AI图像的快速发展模糊了真实与合成图像的界限,引发了生成器与检测器之间的军备竞赛。本文研究了检测器在此竞争中处于劣势的条件,分析了数据维度和数据复杂性两个关键因素。尽管维度提升通常增强检测器识别细微差异的能力,但复杂性的影响更为复杂。通过使用柯尔莫哥洛夫复杂度衡量数据内在结构,我们发现:极端简单或高度复杂的数据集都会降低合成图像的可检测性——生成器几乎完美学习简单数据,而极端多样性则掩盖了生成缺陷;相反,中等复杂度的数据集最有利于检测,因为生成器无法完全捕捉数据分布,其错误仍可被察觉。
原文摘要 · Abstract (English)
The rapid progress of image generative AI has blurred the boundary between synthetic and real images, fueling an arms race between generators and discriminators. This paper investigates the conditions under which discriminators are most disadvantaged in this competition. We analyze two key factors: data dimensionality and data complexity. While increased dimensionality often strengthens the discriminators ability to detect subtle inconsistencies, complexity introduces a more nuanced effect. Using Kolmogorov complexity as a measure of intrinsic dataset structure, we show that both very simple and highly complex datasets reduce the detectability of synthetic images; generators can learn simple datasets almost perfectly, whereas extreme diversity masks imperfections. In contrast, intermediate-complexity datasets create the most favorable conditions for detection, as generators fail to fully capture the distribution and their errors remain visible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。