构建首个涵盖生成与真实内容的多模态健康伪信息数据集
From Generation to Detection: A Multimodal Multi-Task Dataset for Benchmarking Health Misinformation
- 收集34,746篇含图文的健康类信息,区分人类与AI生成
- 现有模型在识别信息真伪和来源上准确率不足,暴露检测短板
- 适合研究虚假信息检测、AI生成内容识别的研究者使用
信息疫情与健康伪信息对个人和社会造成严重负面影响,加剧了公众对推荐健康措施的犹豫。生成式AI能够产出逼真的人类风格文本与图像,显著加速并扩大了健康伪信息的传播,导致其扩散量急剧上升。现有研究多基于社交媒体和事实核查平台构建伪信息数据集,但存在主题覆盖有限、缺乏AI生成内容、原始数据难以获取等问题。为此,我们提出MM Health,一个大规模多模态健康伪信息数据集,包含34,746篇新闻文章,涵盖文本与视觉信息。其中5,776篇为人类生成,28,880篇由多种主流生成式AI模型产生。我们针对可靠性判断、原创性检查及细粒度AI生成检测三个任务进行了基准测试,结果显示现有先进模型在区分信息可靠性和来源方面表现不佳。该数据集旨在支持多场景下健康伪信息的检测,推动跨模态的人类与机器生成内容识别技术发展。
原文摘要 · Abstract (English)
Infodemics and health misinformation have significant negative impact on individuals and society, exacerbating confusion and increasing hesitancy in adopting recommended health measures. Recent advancements in generative AI, capable of producing realistic, human like text and images, have significantly accelerated the spread and expanded the reach of health misinformation, resulting in an alarming surge in its dissemination. To combat the infodemics, most existing work has focused on developing misinformation datasets from social media and fact checking platforms, but has faced limitations in topical coverage, inclusion of AI generation, and accessibility of raw content. To address these issues, we present MM Health, a large scale multimodal misinformation dataset in the health domain consisting of 34,746 news article encompassing both textual and visual information. MM Health includes human-generated multimodal information (5,776 articles) and AI generated multimodal information (28,880 articles) from various SOTA generative AI models. Additionally, We benchmarked our dataset against three tasks (reliability checks, originality checks, and fine-grained AI detection) demonstrating that existing SOTA models struggle to accurately distinguish the reliability and origin of information. Our dataset aims to support the development of misinformation detection across various health scenarios, facilitating the detection of human and machine generated content at multimodal levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。