arXiv:2508.09999cs.CLcs.LG2025-08被引 6

构建真实世界多模态假信息数据集,助力大模型检测能力评估

XFacta: Contemporary, Real-World Dataset and Evaluation for Multimodal Misinformation Detection with Multimodal LLMs

  • 基于真实社交媒体事件构建动态数据集,避免过时或虚构内容
  • 系统评测多种大模型架构在假信息检测中的表现,发现关键瓶颈
  • 支持持续更新的检测闭环框架,适合研究者和平台方参考

社交媒体上多模态假信息快速传播,亟需更有效、鲁棒的检测方法。近年来,多模态大语言模型(MLLMs)展现出解决该问题的潜力。然而,现有方法的瓶颈仍不明确(证据检索 vs 推理),阻碍了领域进展。现有基准要么包含过时事件,导致模型因记忆而产生偏差;要么为人工合成,无法反映真实假信息模式。此外,缺乏对MLLM检测策略的全面分析。为此,我们提出XFacta——一个当代、真实世界的数据集,更适于评估基于MLLM的检测器。我们系统评估了不同MLLM架构与规模下的检测策略,并与现有方法进行对比。基于分析结果,进一步构建半自动检测-闭环框架,持续更新数据以保持时效性。研究为推进多模态假信息检测提供了重要洞见与实践指导。代码与数据已公开。

原文摘要 · Abstract (English)

The rapid spread of multimodal misinformation on social media calls for more effective and robust detection methods. Recent advances leveraging multimodal large language models (MLLMs) have shown the potential in addressing this challenge. However, it remains unclear exactly where the bottleneck of existing approaches lies (evidence retrieval v.s. reasoning), hindering the further advances in this field. On the dataset side, existing benchmarks either contain outdated events, leading to evaluation bias due to discrepancies with contemporary social media scenarios as MLLMs can simply memorize these events, or artificially synthetic, failing to reflect real-world misinformation patterns. Additionally, it lacks comprehensive analyses of MLLM-based model design strategies. To address these issues, we introduce XFacta, a contemporary, real-world dataset that is better suited for evaluating MLLM-based detectors. We systematically evaluate various MLLM-based misinformation detection strategies, assessing models across different architectures and scales, as well as benchmarking against existing detection methods. Building on these analyses, we further enable a semi-automatic detection-in-the-loop framework that continuously updates XFacta with new content to maintain its contemporary relevance. Our analysis provides valuable insights and practices for advancing the field of multimodal misinformation detection. The code and data have been released.

假信息检测多模态大模型数据集社交媒体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。