用多阶段推理增强视觉文本一致性验证,提升新闻真伪判断能力
ContextGuard-LVLM: Enhancing News Veracity through Fine-grained Cross-modal Contextual Consistency Verification
- 基于视觉语言大模型构建多阶段推理框架
- 在复杂逻辑与细微语境上显著优于现有方法
- 适合反虚假新闻、媒体审核等场景使用
数字新闻的泛滥亟需可靠的真伪验证方法,尤其关注图文之间深层语境的一致性。传统方法难以解决细粒度跨模态语境一致性(FCCC)问题,即不仅限于实体匹配,还需对视觉叙事、情感基调和背景信息进行深入对齐。为此,我们提出ContextGuard-LVLM框架,基于先进视觉语言大模型(LVLMs),融合多阶段上下文推理机制,并通过强化或对抗学习增强模型,可识别零样本基线遗漏的细微语境错配。我们在三个现有数据集(TamperedNews-Ent、News400-Ent、MMG-Ent)基础上,新增细粒度上下文标注,包括‘上下文情感’、‘视觉叙事主题’和‘场景事件逻辑连贯性’,并引入综合的CTXT(上下文连贯性)实体类型。大量实验表明,ContextGuard-LVLM在几乎所有细粒度一致性任务中均优于当前最先进的零样本LVLM基线(InstructBLIP 和 LLaVA 1.5),在复杂逻辑推理与细微语境理解上表现突出。此外,该模型对微小扰动更具鲁棒性,且在挑战性样本上与人类专家判断一致率更高,验证了其在识别复杂语境脱节方面的有效性。
原文摘要 · Abstract (English)
The proliferation of digital news media necessitates robust methods for verifying content veracity, particularly regarding the consistency between visual and textual information. Traditional approaches often fall short in addressing the fine-grained cross-modal contextual consistency (FCCC) problem, which encompasses deeper alignment of visual narrative, emotional tone, and background information with text, beyond mere entity matching. To address this, we propose ContextGuard-LVLM, a novel framework built upon advanced Vision-Language Large Models (LVLMs) and integrating a multi-stage contextual reasoning mechanism. Our model is uniquely enhanced through reinforced or adversarial learning paradigms, enabling it to detect subtle contextual misalignments that evade zero-shot baselines. We extend and augment three established datasets (TamperedNews-Ent, News400-Ent, MMG-Ent) with new fine-grained contextual annotations, including "contextual sentiment," "visual narrative theme," and "scene-event logical coherence," and introduce a comprehensive CTXT (Contextual Coherence) entity type. Extensive experiments demonstrate that ContextGuard-LVLM consistently outperforms state-of-the-art zero-shot LVLM baselines (InstructBLIP and LLaVA 1.5) across nearly all fine-grained consistency tasks, showing significant improvements in complex logical reasoning and nuanced contextual understanding. Furthermore, our model exhibits superior robustness to subtle perturbations and a higher agreement rate with human expert judgments on challenging samples, affirming its efficacy in discerning sophisticated forms of context detachment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。