arXiv:2603.26052cs.CVcs.AI2026-03中稿 · CVPR

通过像素与文字的局部语义融合,精准识别多模态虚假信息

Bridging Pixels and Words: Mask-Aware Local Semantic Fusion for Multimodal Media Verification

  • 用掩码标签对作为锚点,实现图文跨模态主动验证
  • 双流查询机制可显式定位图文间细微语义矛盾
  • 适合需要高精度反虚假信息的检测场景

随着多模态虚假信息日益复杂,其检测与定位变得至关重要。然而,现有方法依赖被动的整体融合,难以应对复杂虚假内容。由于‘特征稀释’问题,全局对齐常会平均掉细微的局部语义不一致,反而掩盖了本应发现的矛盾。我们提出 MaLSF(Mask-aware Local Semantic Fusion)框架,将验证范式转向主动、双向交叉验证,模拟人类认知中的交叉比对。MaLSF 利用掩码-标签对作为语义锚点,连接像素与文本。核心机制包含两项创新:1)双向跨模态验证(BCV)模块,作为提问者,使用并行查询流(文本为查询和图像为查询),显式定位冲突;2)分层语义聚合(HSA)模块,智能整合多粒度冲突信号以支持任务特定推理。此外,为提取细粒度掩码-标签对,我们引入一组多样化的解析器。MaLSF 在 DGM4 与多模态假新闻检测任务上均达到当前最优性能。大量消融实验与可视化结果进一步验证了其有效性与可解释性。

原文摘要 · Abstract (English)

As multimodal misinformation becomes more sophisticated, its detection and grounding are crucial. However, current multimodal verification methods, relying on passive holistic fusion, struggle with sophisticated misinformation. Due to 'feature dilution,' global alignments tend to average out subtle local semantic inconsistencies, effectively masking the very conflicts they are designed to find. We introduce MaLSF (Mask-aware Local Semantic Fusion), a novel framework that shifts the paradigm to active, bidirectional verification, mimicking human cognitive cross-referencing. MaLSF utilizes mask-label pairs as semantic anchors to bridge pixels and words. Its core mechanism features two innovations: 1) a Bidirectional Cross-modal Verification (BCV) module that acts as an interrogator, using parallel query streams (Text-as-Query and Image-as-Query) to explicitly pinpoint conflicts; and 2) a Hierarchical Semantic Aggregation (HSA) module that intelligently aggregates these multi-granularity conflict signals for task-specific reasoning. In addition, to extract fine-grained mask-label pairs, we introduce a set of diverse mask-label pair extraction parsers. MaLSF achieves state-of-the-art performance on both the DGM4 and multimodal fake news detection tasks. Extensive ablation studies and visualization results further verify its effectiveness and interpretability.

多模态验证虚假信息检测语义融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。