提升视觉语言模型对图文错位谣言的判断力与解释力
E2LVLM:Evidence-Enhanced Large Vision-Language Model for Multimodal Out-of-Context Misinformation Detection
- 用重排重写策略让检索文本更贴合模型输入
- 在新闻领域构建新数据集,支持判断与解释同步训练
- 不仅识别谣言,还能给出可信推理过程,适合需要透明决策场景
近期大型多模态视觉语言模型(LVLM)在图文错位谣言检测上取得显著进展,可判断真实图像是否被错误引用。然而,现有方法直接将反向搜索获取的文本证据输入模型,导致决策阶段出现误判。为此,本文提出E2LVLM,通过两级证据增强机制改进:首先,针对外部工具提供的文本证据与模型输入不匹配的问题,设计重排与重写策略,生成更连贯、语境契合的内容,提升模型对真实图像的理解能力;其次,为解决新闻领域缺乏带判断与解释标注的数据集问题,利用提示工程引导LVLM生成合理解释,构建新型多模态指令跟随数据集;并采用包含可信解释的多模态指令微调策略,实现超越单纯检测的可解释判断。大量实验证明,E2LVLM在性能上优于现有最佳方法,并能提供有力推理依据。
原文摘要 · Abstract (English)
Recent studies in Large Vision-Language Models (LVLMs) have demonstrated impressive advancements in multimodal Out-of-Context (OOC) misinformation detection, discerning whether an authentic image is wrongly used in a claim. Despite their success, the textual evidence of authentic images retrieved from the inverse search is directly transmitted to LVLMs, leading to inaccurate or false information in the decision-making phase. To this end, we present E2LVLM, a novel evidence-enhanced large vision-language model by adapting textual evidence in two levels. First, motivated by the fact that textual evidence provided by external tools struggles to align with LVLMs inputs, we devise a reranking and rewriting strategy for generating coherent and contextually attuned content, thereby driving the aligned and effective behavior of LVLMs pertinent to authentic images. Second, to address the scarcity of news domain datasets with both judgment and explanation, we generate a novel OOC multimodal instruction-following dataset by prompting LVLMs with informative content to acquire plausible explanations. Further, we develop a multimodal instruction-tuning strategy with convincing explanations for beyond detection. This scheme contributes to E2LVLM for multimodal OOC misinformation detection and explanation. A multitude of experiments demonstrate that E2LVLM achieves superior performance than state-of-the-art methods, and also provides compelling rationales for judgments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。