arXiv:2505.20122cs.CV2025-05被引 1

提出新基准MEBench,评估视觉语言模型的互斥性偏差。

MEBench: A Novel Benchmark for Understanding Mutual Exclusivity Bias in Vision-Language Models

  • 引入空间推理增强互斥性测试场景。
  • 模型在多对象情境中仍显弱互斥性,但能利用空间信息解歧。
  • 适合研究认知偏差与视觉语言理解的学者。

本文提出MEBench,一个用于评估视觉语言模型中互斥性(ME)偏差的新基准。该基准借鉴儿童词义学习中的认知现象,通过引入空间推理,构建更具挑战性和真实性的评估环境。为支持可控实验,我们设计了灵活可扩展的数据生成管道,可生成多样化的带标注场景。采用新颖的评估指标衡量多种视觉语言模型在该基准上的表现,发现这些模型表现出较弱的互斥性偏差,但在存在额外空间上下文时,能部分利用其解决多个新物体带来的语义模糊问题。项目页面:http://mebench.github.io/。

原文摘要 · Abstract (English)

This paper introduces MEBench, a novel benchmark for evaluating mutual exclusivity (ME) bias, a cognitive phenomenon observed in children during word learning. Unlike traditional ME tasks, MEBench further incorporates spatial reasoning to create more challenging and realistic evaluation settings. To facilitate controlled experimentation, we also present a flexible and scalable data generation pipeline that supports the construction of diverse annotated scenes. We assess the performance of various vision-language models (VLMs) on this benchmark using novel evaluation metrics that capture key aspects of ME-based reasoning. We find that these VLMs exhibit weak ME bias, while showing some ability to leverage extra spatial context to resolve ambiguity in multiple novel object settings. Project page: http://mebench.github.io/.

视觉语言模型认知偏差评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。