arXiv:2508.20670cs.CVcs.MM2025-08被引 3

构建首个面向生成图像意图的多模态数据集,区分幽默、艺术与虚假信息。

"Humor, Art, or Misinformation?": A Multimodal Dataset for Intent-Aware Synthetic Image Detection

  • 基于社交媒体真实图文对,标注三类生成图像意图。
  • 多模态引导生成数据使模型在真实场景下泛化能力更强。
  • 揭示当前模型对意图理解仍有限,需专门架构支持。

近期多模态人工智能在检测合成与语境错位内容方面取得进展,但现有研究普遍忽视生成图像背后的意图。为填补这一空白,我们提出 S-HArM,一个用于意图感知分类的多模态数据集,包含来自 Twitter/X 与 Reddit 的 9,576 组“真实场景”图像-文本对,标注为幽默/讽刺、艺术或虚假信息。此外,我们探索了三种提示策略(图像引导、描述引导、多模态引导),利用 Stable Diffusion 构建大规模合成训练数据集。通过对比多种方法,包括模态融合、对比学习、重建网络、注意力机制及大型视觉-语言模型,结果表明:在图像和多模态引导数据上训练的模型能更好地泛化至真实场景内容,因其保留了视觉上下文。然而,整体性能仍受限,凸显推断意图的复杂性,亟需专用模型架构。

原文摘要 · Abstract (English)

Recent advances in multimodal AI have enabled progress in detecting synthetic and out-of-context content. However, existing efforts largely overlook the intent behind AI-generated images. To fill this gap, we introduce S-HArM, a multimodal dataset for intent-aware classification, comprising 9,576 "in the wild" image-text pairs from Twitter/X and Reddit, labeled as Humor/Satire, Art, or Misinformation. Additionally, we explore three prompting strategies (image-guided, description-guided, and multimodally-guided) to construct a large-scale synthetic training dataset with Stable Diffusion. We conduct an extensive comparative study including modality fusion, contrastive learning, reconstruction networks, attention mechanisms, and large vision-language models. Our results show that models trained on image- and multimodally-guided data generalize better to "in the wild" content, due to preserved visual context. However, overall performance remains limited, highlighting the complexity of inferring intent and the need for specialized architectures.

图像检测多模态意图识别生成内容

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。