arXiv:2607.24016cs.CV2026-07

新基准DailyBench测试生成与编辑图像的检测能力,发现现有方法严重失效。

DailyBench: A Unified Benchmark for AI-Generated and Manipulated Images from Modern Generative Models

论文配图:DailyBench: A Unified Benchmark for AI-Generated and Manipulated Images from Modern Generative Models
图 1 · 摘自论文原文
  • 构建合成与局部编辑双子集,模拟真实生成与修图场景。
  • 现有检测器在真实数据上准确率暴跌至54%-76%,暴露泛化缺陷。
  • 适合研究鲁棒检测与抗篡改算法的学者使用。

近期生成模型的进步使图像检测从识别明显合成图像转向应对高度真实的生成与编辑内容。然而,现有检测基准多基于过时模型,侧重全图合成,与现实中的生成和编辑场景严重脱节。为此,我们提出DailyBench——一个高质量统一基准,用于评估检测器在现代全图生成与对象级编辑下的泛化能力。该基准包含两个互补子集:FakeBench涵盖由最新开源与商用生成模型产生的高质合成图像;ManipulationBench引入先进图像条件模型对真实图像进行挑战性对象级编辑。这一设计使DailyBench成为研究生成器泛化与感知局部编辑的现实测试平台。实验表明,当前检测器在GenImage上报告91-96%平衡准确率,在FakeBench上骤降至60-76%,在ManipulationBench上进一步降至54-66%。结果表明现有检测器对真实生成与编辑仍缺乏泛化能力,凸显DailyBench作为严谨测试平台的价值。项目地址:https://dailybench.github.io/

原文摘要 · Abstract (English)

Recent advances in generative models have shifted AI-generated image detection from identifying easily distinguishable, fully synthetic images to identifying highly realistic content generated by both modern generation and manipulation pipelines. However, existing detection benchmarks are often built with outdated generative models and primarily emphasize full-image synthesis, creating a growing mismatch between benchmark data and the images encountered in real-world generation and editing scenarios. To bridge this gap, we introduce DailyBench, a high-quality unified benchmark for evaluating whether AI-generated image detectors can generalize across both modern full-image synthesis and object-level manipulation. DailyBench contains two complementary subsets: FakeBench, which includes high-quality images synthesized by recent open-source and commercial generative models, and ManipulationBench, which introduces challenging object-level edits applied to real images using advanced image-conditional models. This design makes DailyBench a realistic testbed for studying both generator-level generalization and manipulation-aware detection under subtle local edits. Experiments on DailyBench reveal substantial robustness gaps in current detectors: methods reporting 91-96% balanced accuracy on GenImage drop to 60-76% on FakeBench and 54-66% on ManipulationBench. These results show that existing detectors remain poorly generalized to realistic synthesis and manipulation, highlighting DailyBench as a rigorous testbed for developing robust and manipulation-aware AI-generated image detection methods. The project is available at https://dailybench.github.io/

图像检测生成模型真实性评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。