arXiv:2512.20257cs.CV2025-12

轻量级多模态假信息检测模型,少标注也能高效准确。

LADLE-MM: Limited Annotation based Detector with Learned Ensembles for Multimodal Misinformation

  • 用模型集成+固定参考空间,减少对标注数据依赖。
  • 参数少60.3%仍达顶尖性能,无地标的任务上表现更优。
  • 适合资源有限但需抗单模态偏见的假信息检测场景。

随着多媒体内容生成与编辑工具的普及,跨模态的合成篡改已成严重威胁,常被用于扭曲重要事件叙述并在社交媒体传播假信息。针对图像-文本对形式的假信息检测,现有方法通常依赖计算量大的架构或大量标注数据。本文提出LADLE-MM:一种基于模型集成的轻量级多模态假信息检测器,专为标注有限、训练资源受限的场景设计。该模型包含两个单模态分支和一个融合模态分支,通过BLIP提取的固定多模态嵌入增强图像与文本表征。相比之前最先进模型,其可训练参数减少60.3%,在DGM4基准上于二分类与多标签任务中均表现优异,且在无地标的训练条件下超越现有方法。在VERITE数据集上的评估显示,尽管未使用大型视觉语言模型,其仍优于复杂架构的当前最优方法,展现出强泛化能力与对单模态偏见的鲁棒性。

原文摘要 · Abstract (English)

With the rise of easily accessible tools for generating and manipulating multimedia content, realistic synthetic alterations to digital media have become a widespread threat, often involving manipulations across multiple modalities simultaneously. Recently, such techniques have been increasingly employed to distort narratives of important events and to spread misinformation on social media, prompting the development of misinformation detectors. In the context of misinformation conveyed through image-text pairs, several detection methods have been proposed. However, these approaches typically rely on computationally intensive architectures or require large amounts of annotated data. In this work we introduce LADLE-MM: Limited Annotation based Detector with Learned Ensembles for Multimodal Misinformation, a model-soup initialized multimodal misinformation detector designed to operate under a limited annotation setup and constrained training resources. LADLE-MM is composed of two unimodal branches and a third multimodal one that enhances image and text representations with additional multimodal embeddings extracted from BLIP, serving as fixed reference space. Despite using 60.3% fewer trainable parameters than previous state-of-the-art models, LADLE-MM achieves competitive performance on both binary and multi-label classification tasks on the DGM4 benchmark, outperforming existing methods when trained without grounding annotations. Moreover, when evaluated on the VERITE dataset, LADLE-MM outperforms current state-of-the-art approaches that utilize more complex architectures involving Large Vision-Language-Models, demonstrating the effective generalization ability in an open-set setting and strong robustness to unimodal bias.

假信息检测多模态轻量化少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。