arXiv:2504.20658cs.MMcs.AI2025-04被引 12

构建60万张真实社交网络传播的假图像数据集,用于测试检测模型在现实场景下的表现。

TrueFake: A Real World Case Dataset of Last Generation Fake Images also Shared on Social Networks

  • 收集60万张通过三大社交平台传播的顶级生成假图像
  • 发现社交媒体压缩会显著降低现有检测模型准确率
  • 为真实世界假图像检测提供可复现的基准测试

AI生成的合成媒体正被广泛用于现实场景,常以在社交平台上散播虚假信息和宣传为目的,而压缩等处理会破坏假图像检测线索。当前许多取证工具未考虑这些真实环境挑战。本文提出TrueFake,一个包含60万张图像的大规模基准数据集,涵盖顶尖生成技术及通过三个不同社交网络的传播。该数据集支持在高度真实且具有挑战性的条件下,对最先进的假图像检测器进行严格评估。通过大量实验,我们分析了社交网络传播对检测性能的影响,并识别出当前最有效的检测与训练策略。研究强调必须在贴近真实使用场景的条件下评估取证模型。

原文摘要 · Abstract (English)

AI-generated synthetic media are increasingly used in real-world scenarios, often with the purpose of spreading misinformation and propaganda through social media platforms, where compression and other processing can degrade fake detection cues. Currently, many forensic tools fail to account for these in-the-wild challenges. In this work, we introduce TrueFake, a large-scale benchmarking dataset of 600,000 images including top notch generative techniques and sharing via three different social networks. This dataset allows for rigorous evaluation of state-of-the-art fake image detectors under very realistic and challenging conditions. Through extensive experimentation, we analyze how social media sharing impacts detection performance, and identify current most effective detection and training strategies. Our findings highlight the need for evaluating forensic models in conditions that mirror real-world use.

假图像检测真实场景数据集社交网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。