arXiv:2601.07178cs.CVcs.AI2026-01被引 1

动态迭代视觉证据推理,提升多模态假新闻检测准确率与效率

DIVER: Dynamic Iterative Visual Evidence Reasoning for Multimodal Fake News Detection

  • 基于文本先行、视觉按需引入的渐进式推理框架
  • 在三个数据集上平均比最优基线高2.72%准确率,延迟降低4.12秒
  • 适合需要高精度与低延迟假新闻检测的系统开发者

多模态假新闻检测对遏制恶意信息传播至关重要。现有方法依赖静态融合或大语言模型,因视觉基础薄弱而存在计算冗余和幻觉风险。为此,我们提出DIVER(动态迭代视觉证据推理)框架,采用渐进式、证据驱动的推理范式。DIVER首先通过语言分析建立强文本基线,利用模内一致性过滤不可靠或幻觉性陈述;仅当文本证据不足时,才引入视觉信息,并通过跨模态对齐验证自适应判断是否需深入视觉分析。对于跨模态语义差异显著的样本,DIVER选择性调用细粒度视觉工具(如OCR、密集描述),通过不确定度感知融合机制迭代聚合任务相关证据,以优化多模态推理。在Weibo、Weibo21和GossipCop数据集上的实验表明,DIVER平均优于当前最佳基线2.72%,同时将推理延迟降低4.12秒。

原文摘要 · Abstract (English)

Multimodal fake news detection is crucial for mitigating adversarial misinformation. Existing methods, relying on static fusion or LLMs, face computational redundancy and hallucination risks due to weak visual foundations. To address this, we propose DIVER (Dynamic Iterative Visual Evidence Reasoning), a framework grounded in a progressive, evidence-driven reasoning paradigm. DIVER first establishes a strong text-based baseline through language analysis, leveraging intra-modal consistency to filter unreliable or hallucinated claims. Only when textual evidence is insufficient does the framework introduce visual information, where inter-modal alignment verification adaptively determines whether deeper visual inspection is necessary. For samples exhibiting significant cross-modal semantic discrepancies, DIVER selectively invokes fine-grained visual tools (e.g., OCR and dense captioning) to extract task-relevant evidence, which is iteratively aggregated via uncertainty-aware fusion to refine multimodal reasoning. Experiments on Weibo, Weibo21, and GossipCop demonstrate that DIVER outperforms state-of-the-art baselines by an average of 2.72\%, while optimizing inference efficiency with a reduced latency of 4.12 s.

假新闻检测多模态推理视觉证据高效检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。