arXiv:2601.15316cs.AIcs.CV2026-01综述被引 10

用大模型统一分析图文,提升假新闻识别能力

The Paradigm Shift: A Comprehensive Survey on Large Vision Language Models for Multimodal Fake News Detection

  • 以大视觉语言模型实现图文端到端联合推理
  • 突破传统方法在语义理解上的局限
  • 适合关注多媒体内容安全的研究者

近年来,大视觉语言模型(LVLMs)的快速发展推动了多模态假新闻检测(MFND)的范式变革,将检测方法从传统的特征工程转向统一的端到端多模态推理框架。早期方法主要依赖浅层融合技术捕捉文本与图像间的关联,但难以实现高层语义理解与复杂的跨模态交互。LVLMs的出现通过强大的表征学习能力实现了视觉与语言的联合建模,显著提升了对同时利用文字叙述和视觉内容的虚假信息的识别能力。尽管如此,该领域仍缺乏系统性综述来梳理这一演进过程并整合最新进展。本文首次全面回顾了基于LVLM的MFND研究,首先呈现从传统多模态检测流程到基础模型驱动范式的演进历程;其次建立涵盖模型架构、数据集与性能基准的结构化分类体系;进一步分析当前关键技术挑战,包括可解释性、时间推理与领域泛化问题;最后展望未来研究方向,以指导该范式变革的下一阶段发展。现有方法总结见:https://github.com/Tan-YiLong/Overview-of-Fake-News-Detection。

原文摘要 · Abstract (English)

In recent years, the rapid evolution of large vision-language models (LVLMs) has driven a paradigm shift in multimodal fake news detection (MFND), transforming it from traditional feature-engineering approaches to unified, end-to-end multimodal reasoning frameworks. Early methods primarily relied on shallow fusion techniques to capture correlations between text and images, but they struggled with high-level semantic understanding and complex cross-modal interactions. The emergence of LVLMs has fundamentally changed this landscape by enabling joint modeling of vision and language with powerful representation learning, thereby enhancing the ability to detect misinformation that leverages both textual narratives and visual content. Despite these advances, the field lacks a systematic survey that traces this transition and consolidates recent developments. To address this gap, this paper provides a comprehensive review of MFND through the lens of LVLMs. We first present a historical perspective, mapping the evolution from conventional multimodal detection pipelines to foundation model-driven paradigms. Next, we establish a structured taxonomy covering model architectures, datasets, and performance benchmarks. Furthermore, we analyze the remaining technical challenges, including interpretability, temporal reasoning, and domain generalization. Finally, we outline future research directions to guide the next stage of this paradigm shift. To the best of our knowledge, this is the first comprehensive survey to systematically document and analyze the transformative role of LVLMs in combating multimodal fake news. The summary of existing methods mentioned is in our Github: \href{https://github.com/Tan-YiLong/Overview-of-Fake-News-Detection}{https://github.com/Tan-YiLong/Overview-of-Fake-News-Detection}.

假新闻检测视觉语言模型多模态综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。