arXiv:2508.11317cs.CVcs.MM2025-08AAAI被引 8

发现视觉语言模型的逻辑盲区,提出新训练方法提升推理能力

Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models

  • 构建涵盖9类逻辑的5万+数据集,系统评估模型逻辑能力
  • 现有模型在因果/条件推理上比人类低40分以上,依赖表面语义
  • 提出LogicCLIP框架,兼顾逻辑敏感性与通用对齐性能

视觉语言模型(VLMs)如CLIP已成为多模态智能的基础,但其逻辑理解能力尚未充分探索,导致实际应用中存在关键的‘逻辑盲区’。为系统诊断该问题,我们提出LogicBench,一个包含超过5万组视觉-语言对的综合性基准,覆盖9类逻辑范畴及4种场景:图像、视频、异常检测与医学诊断。评估显示,现有VLMs即使是最先进的模型,在因果性和条件性等挑战任务上也低于人类表现40分以上,表明其依赖表面语义而非深层逻辑结构。为弥补这一差距,我们提出LogicCLIP,一种新型训练框架,通过逻辑感知的数据生成和优化目标设计,提升模型逻辑敏感性。该方法结合粗粒度对齐、细粒度多选目标与新颖的逻辑结构感知目标。大量实验表明,LogicCLIP在所有LogicBench领域显著优于基线,且保持甚至超越通用视觉-语言基准性能,证明逻辑增强不以牺牲通用对齐为代价。我们相信LogicBench与LogicCLIP将推动VLM逻辑能力的发展。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs), exemplified by CLIP, have emerged as foundational for multimodal intelligence. However, their capacity for logical understanding remains significantly underexplored, resulting in critical ''logical blindspots'' that limit their reliability in practical applications. To systematically diagnose this, we introduce LogicBench, a comprehensive benchmark with over 50,000 vision-language pairs across 9 logical categories and 4 diverse scenarios: images, videos, anomaly detection, and medical diagnostics. Our evaluation reveals that existing VLMs, even the state-of-the-art ones, fall at over 40 accuracy points below human performance, particularly in challenging tasks like Causality and Conditionality, highlighting their reliance on surface semantics over critical logical structures. To bridge this gap, we propose LogicCLIP, a novel training framework designed to boost VLMs' logical sensitivity through advancements in both data generation and optimization objectives. LogicCLIP utilizes logic-aware data generation and a contrastive learning strategy that combines coarse-grained alignment, a fine-grained multiple-choice objective, and a novel logical structure-aware objective. Extensive experiments demonstrate LogicCLIP's substantial improvements in logical comprehension across all LogicBench domains, significantly outperforming baselines. Moreover, LogicCLIP retains, and often surpasses, competitive performance on general vision-language benchmarks, demonstrating that the enhanced logical understanding does not come at the expense of general alignment. We believe that LogicBench and LogicCLIP will be important resources for advancing VLM logical capabilities.

视觉语言模型逻辑推理基准测试多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。