arXiv:2508.10444cs.CL2025-08被引 3

用多样真实相关的推理提升多模态假信息检测效果

DiFaR: Enhancing Multimodal Misinformation Detection with Diverse, Factual, and Relevant Rationales

  • 设计五种思维链提示+轻量过滤模块,生成高质量推理
  • 在四个数据集上最高提升检测准确率8.7%
  • 适合需要可解释假信息检测的开发者和研究者

从大型视觉语言模型(LVLMs)生成文本推理以支持可训练的多模态假信息检测已成为一种有前景的方法。然而,其有效性受到三个核心挑战的根本限制:(i) 推理内容多样性不足,(ii) 因幻觉导致的事实错误,(iii) 无关或矛盾内容引入噪声。我们提出DiFaR,一个检测器无关的框架,用于生成多样、真实且相关的推理以增强假信息检测。DiFaR采用五种思维链提示从LVLMs中激发多样的推理路径,并引入一个轻量级后处理过滤模块,基于句子级事实性和相关性得分筛选推理语句。在四个主流基准上的大量实验表明,DiFaR相较于四类基线方法最高提升5.9%,并使现有检测器性能提升达8.7%。自动指标与人工评估均证实,DiFaR在所有三个维度上显著提升了推理质量。

原文摘要 · Abstract (English)

Generating textual rationales from large vision-language models (LVLMs) to support trainable multimodal misinformation detectors has emerged as a promising paradigm. However, its effectiveness is fundamentally limited by three core challenges: (i) insufficient diversity in generated rationales, (ii) factual inaccuracies due to hallucinations, and (iii) irrelevant or conflicting content that introduces noise. We introduce DiFaR, a detector-agnostic framework that produces diverse, factual, and relevant rationales to enhance misinformation detection. DiFaR employs five chain-of-thought prompts to elicit varied reasoning traces from LVLMs and incorporates a lightweight post-hoc filtering module to select rationale sentences based on sentence-level factuality and relevance scores. Extensive experiments on four popular benchmarks demonstrate that DiFaR outperforms four baseline categories by up to 5.9% and boosts existing detectors by as much as 8.7%. Both automatic metrics and human evaluations confirm that DiFaR significantly improves rationale quality across all three dimensions.

假信息检测视觉语言模型可解释性推理生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。