构建新基准,揭示视觉模型跨模态失败模式。
AMVICC: A Novel Benchmark for Cross-Modal Failure Mode Profiling for VLMs and IGMs
- 设计AMVICC基准,对比图文任务中的失败模式。
- 11个MLLM和3个IGM在9类任务中表现共享失败特征。
- 生成模型对细粒度视觉属性控制差,尤其在显式提示下。
我们通过创建一个新基准,系统比较图像到文本与文本到图像任务中的失败模式,实现对视觉理解的跨模态评估。尽管机器学习快速发展,视觉语言模型(VLMs)仍难以理解基本视觉概念如物体朝向、数量和空间关系,暴露出基础视觉推理的不足。通过将MMVP基准问题转化为显式和隐式提示,我们构建了AMVICC,用于跨模态失败模式分析。测试11个MLLM和3个IGM在9类视觉推理任务后发现,失败模式常在模型与模态间共享,但部分失败具模型或模态特异性,可能源于多种因素。IGMs在响应提示时普遍难以操控特定视觉成分,尤其在显式提示下,表明其对细粒度视觉属性控制能力弱。研究结果直接适用于当前先进模型在结构化视觉推理任务中的评估,为未来跨模态对齐研究提供框架,帮助判断图像生成与视觉理解失败是否源自共同局限性。这些洞察可指导统一视觉语言建模的改进。
原文摘要 · Abstract (English)
We investigate visual reasoning limitations of both multimodal large language models (MLLMs) and image generation models (IGMs) by creating a novel benchmark to systematically compare failure modes across image-to-text and text-to-image tasks, enabling cross-modal evaluation of visual understanding. Despite rapid growth in machine learning, vision language models (VLMs) still fail to understand basic visual concepts such as object orientation, quantity, and spatial relationships, which highlights gaps in elementary visual reasoning. By adapting MMVP benchmark questions into explicit and implicit prompts, we create \textit{AMVICC}, a novel benchmark for profiling failure modes across various modalities. After testing 11 MLLMs and 3 IGMs in 9 categories of visual reasoning, our results show that failure modes are often shared between models and modalities. However, certain failures are model-specific and modality-specific, and this can potentially be attributed to various factors. IGMs consistently struggle to manipulate specific visual components in response to prompts, especially in explicit prompts, suggesting poor control over fine-grained visual attributes. Our findings apply most directly to the evaluation of existing state-of-the-art models on structured visual reasoning tasks. This work lays the foundation for future cross-modal alignment studies, offering a framework to probe whether image generation and visual interpretation failures stem from shared limitations. These insights can guide future improvements in unified vision-language modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。