测试大模型识别属性修改物体的能力,发现性能仍有明显短板。
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects?
- 构建新基准NEMO,包含900张原始与属性修改的水果图像。
- 26个模型在该基准上表现参差,最强模型准确率仅72.3%。
- 模型越大不等于越好,过大可能削弱视觉编码器性能。
多模态大语言模型(MLLMs)在视觉理解方面取得显著进展,但其对特定属性修改物体的识别能力仍存疑问。为此,我们研究了MLLMs在从常识到超常识场景下的推理能力。提出新基准NEMO,包含900张原始水果及其对应的属性修改图像,配套2,700个问题,涵盖开放、多选、不可解三种类型。评估26个近期开源及商用模型,结果揭示其在识别属性修改物体上存在明显性能差距,且不同模型答案偏好各异。尽管更强的视觉编码器能提升性能,但MLLMs仍落后于独立视觉编码器。有趣的是,模型规模扩大并未持续带来更好结果,深入分析显示更大的语言模型在微调过程中可能损害视觉编码器。这些发现揭示了当前MLLMs的关键局限,并为开发更通用、鲁棒的多模态模型指明方向。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have made notable advances in visual understanding, yet their abilities to recognize objects modified by specific attributes remain an open question. To address this, we explore MLLMs' reasoning capabilities in object recognition, ranging from commonsense to beyond-commonsense scenarios. We introduce a novel benchmark, NEMO, which comprises 900 images of origiNal fruits and their corresponding attributE-MOdified ones; along with a set of 2,700 questions including open-, multiple-choice-, unsolvable types. We assess 26 recent open-sourced and commercial models using our benchmark. The findings highlight pronounced performance gaps in recognizing objects in NEMO and reveal distinct answer preferences across different models. Although stronger vision encoders improve performance, MLLMs still lag behind standalone vision encoders. Interestingly, scaling up the model size does not consistently yield better outcomes, as deeper analysis reveals that larger LLMs can weaken vision encoders during fine-tuning. These insights shed light on critical limitations in current MLLMs and suggest potential pathways toward developing more versatile and resilient multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。