arXiv:2505.07704cs.CVcs.CL2025-05

用视觉语言模型检测图像是否违背常识,判断真实感。

Through the Looking Glass: Common Sense Consistency Evaluation of Weird Images

  • 用大模型提取图像中的原子事实,再通过轻量分类器判断一致性。
  • 在WHOOPS!和WEIRD数据集上达到新最好性能。
  • 适合研究图像真实性评估与常识理解的学者。

衡量图像的真实性在人工智能研究中是一项复杂任务。例如,沙漠中出现一个拿着吸尘器的男孩就违背了常识。我们提出一种新方法Through the Looking Glass(TLG),利用大视觉语言模型(LVLMs)和基于Transformer的编码器评估图像的常识一致性。通过LVLM从图像中提取原子事实,获得混合准确事实,再对编码后的原子事实进行轻量级微调,构建注意力池化分类器。TLG在WHOOPS!和WEIDD数据集上均取得当前最优表现,同时仅需紧凑的微调组件。

原文摘要 · Abstract (English)

Measuring how real images look is a complex task in artificial intelligence research. For example, an image of a boy with a vacuum cleaner in a desert violates common sense. We introduce a novel method, which we call Through the Looking Glass (TLG), to assess image common sense consistency using Large Vision-Language Models (LVLMs) and Transformer-based encoder. By leveraging LVLMs to extract atomic facts from these images, we obtain a mix of accurate facts. We proceed by fine-tuning a compact attention-pooling classifier over encoded atomic facts. Our TLG has achieved a new state-of-the-art performance on the WHOOPS! and WEIRD datasets while leveraging a compact fine-tuning component.

常识推理图像评估视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。