检验视觉语言模型能否用事实验证推理链条,防止幻觉。
LogicGaze: Benchmarking Causal Consistency in Visual Narratives via Counterfactual Verification
- 构建反事实验证框架,测试模型推理是否匹配真实画面。
- 在4万段视频和部分Flickr30k图像上暴露顶尖模型幻觉问题。
- 适合关注多模态可靠性与可信推理的研究者使用。
尽管顺序推理提升了视觉语言模型(VLMs)执行复杂多模态任务的能力,但其推理链是否基于实际视觉证据仍缺乏充分探究。我们提出LogicGaze,一个新型基准框架,旨在严格检验VLMs能否根据视觉输入验证其因果推理链,聚焦于普遍存在的幻觉问题。数据源自40,000段ShareGPT4Video视频片段及Flickr30k子集图像,通过引入在语言上合理但视觉矛盾的扰动,迫使模型验证每一步推理的真实性。采用三阶段评估协议——因果验证、有根基的故事合成与扰动拒绝——揭示了如Qwen2.5-VL-72B等前沿VLMs存在显著漏洞。LogicGaze倡导更稳健、可信的多模态推理,所有资源已匿名公开于开源仓库。
原文摘要 · Abstract (English)
While sequential reasoning enhances the capability of Vision-Language Models (VLMs) to execute complex multimodal tasks, their reliability in grounding these reasoning chains within actual visual evidence remains insufficiently explored. We introduce LogicGaze, a novel benchmark framework designed to rigorously interrogate whether VLMs can validate sequential causal chains against visual inputs, specifically targeting the pervasive issue of hallucination. Curated from 40,000 video segments from ShareGPT4Video and a subset of Flickr30k imagery, LogicGaze integrates causal sequences with visually contradictory yet linguistically plausible perturbations, compelling models to verify the authenticity of each reasoning step. Our tripartite evaluation protocol - Causal Validation, Grounded Narrative Synthesis, and Perturbation Rejection - exposes significant vulnerabilities in state-of-the-art VLMs such as Qwen2.5-VL-72B. LogicGaze advocates for robust, trustworthy multimodal reasoning, with all resources publicly available in an anonymized repository.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。