构建1.25万对图文数据集,检测由AI生成的假新闻。
MiRAGeNews: Multimodal Realistic AI-Generated News Detection
- 基于顶尖生成器构建真实与AI生成图文对数据集。
- 人类与现有模型在检测上准确率分别仅60%和不足24%。
- 新模型MiRAGe在跨域数据上提升5.1%检测性能,适合安全研究者使用。
近年来,煽动性或误导性‘假新闻’内容日益泛滥。同时,利用AI工具生成逼真图像已变得极为便捷,可呈现任何想象中的场景。将二者结合——即生成虚假新闻内容——具有极强的破坏力。为应对这一威胁,我们提出了MiRAGeNews数据集,包含12,500对高质量的真实与AI生成图文对,来自最先进的生成器。该数据集对人类(F-1达60%)和当前最先进多模态大模型(F-1低于24%)构成严峻挑战。基于此数据集,我们训练了一个多模态检测器MiRAGe,其在跨域图像生成器及新闻出版商的数据上,相比现有基线模型提升了5.1% F-1。代码与数据集已公开,以支持未来对AI生成内容的检测研究。
原文摘要 · Abstract (English)
The proliferation of inflammatory or misleading "fake" news content has become increasingly common in recent years. Simultaneously, it has become easier than ever to use AI tools to generate photorealistic images depicting any scene imaginable. Combining these two -- AI-generated fake news content -- is particularly potent and dangerous. To combat the spread of AI-generated fake news, we propose the MiRAGeNews Dataset, a dataset of 12,500 high-quality real and AI-generated image-caption pairs from state-of-the-art generators. We find that our dataset poses a significant challenge to humans (60% F-1) and state-of-the-art multi-modal LLMs (< 24% F-1). Using our dataset we train a multi-modal detector (MiRAGe) that improves by +5.1% F-1 over state-of-the-art baselines on image-caption pairs from out-of-domain image generators and news publishers. We release our code and data to aid future work on detecting AI-generated content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。