测试视觉大模型推断因果的能力,发现其表现远不如人类。
NL-Eye: Abductive NLI for Images
- 将抽象推理任务迁移到视觉领域,要求模型判断图像间因果关系。
- 350组三元组数据中,模型准确率接近随机水平,人类表现优异。
- 适合研究多模态推理、安全预警系统和内容验证的学者使用。
基于视觉语言模型(VLM)的机器人能否在检测到湿滑地面时发出警示?尽管当前VLM表现出色,但其推断结果与原因的能力仍缺乏深入探索。为此,我们提出NL-Eye基准,用于评估VLM在视觉领域的归纳推理能力。NL-Eye将归纳自然语言推理(NLI)任务拓展至视觉场景,要求模型基于前提图像评估假设图像的合理性,并给出解释。该数据集包含350个精心设计的三元组样本(共1,050张图像),涵盖物理、功能、逻辑、情感、文化及社会六类推理类型。数据构建分两步:先撰写文本描述,再通过文生图模型生成图像,全程依赖人工确保高质量与挑战性。实验表明,现有VLM在NL-Eye上表现极差,准确率接近随机基线,而人类在合理性和解释质量上均显著领先。这揭示了现代VLM在归纳推理方面的明显缺陷。NL-Eye为发展具备鲁棒多模态推理能力的VLM提供了关键进展,适用于事故预防机器人和生成视频验证等真实应用场景。
原文摘要 · Abstract (English)
Will a Visual Language Model (VLM)-based bot warn us about slipping if it detects a wet floor? Recent VLMs have demonstrated impressive capabilities, yet their ability to infer outcomes and causes remains underexplored. To address this, we introduce NL-Eye, a benchmark designed to assess VLMs' visual abductive reasoning skills. NL-Eye adapts the abductive Natural Language Inference (NLI) task to the visual domain, requiring models to evaluate the plausibility of hypothesis images based on a premise image and explain their decisions. NL-Eye consists of 350 carefully curated triplet examples (1,050 images) spanning diverse reasoning categories: physical, functional, logical, emotional, cultural, and social. The data curation process involved two steps - writing textual descriptions and generating images using text-to-image models, both requiring substantial human involvement to ensure high-quality and challenging scenes. Our experiments show that VLMs struggle significantly on NL-Eye, often performing at random baseline levels, while humans excel in both plausibility prediction and explanation quality. This demonstrates a deficiency in the abductive reasoning capabilities of modern VLMs. NL-Eye represents a crucial step toward developing VLMs capable of robust multimodal reasoning for real-world applications, including accident-prevention bots and generated video verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。