评测主流多模态模型对复杂场景组合理解能力,发现仍远逊于人类。
Evaluating Compositional Scene Understanding in Multimodal Generative Models
- 构建新评估框架,测试模型对多对象与关系组合的生成与理解能力。
- 模型在超过5个物体或多重关系的复杂场景中表现明显下降,准确率低于人类。
- 适合关注多模态模型局限性与视觉推理研究的研究者参考。
视觉世界本质上是组合性的,场景由物体及其关系构成。因此,计算机视觉系统必须具备组合性理解能力,才能实现鲁棒且可泛化的场景解析。尽管文本到图像模型(如DALL-E 3)和多模态视觉语言模型(如GPT-4V、GPT-4o、Claude Sonnet 3.5、QWEN2-VL-72B、InternVL2.5-38B)取得显著进展,但其是否能准确生成与理解涉及多个对象及关系的场景仍不明确。本文评估了当前主流模型在组合性视觉处理方面的能力,并与人类参与者进行对比。结果显示,这些模型虽在组合任务上较前代有所提升,但在包含超过5个对象及多重关系的复杂场景中,性能仍显著低于人类,表明组合性理解仍需进一步突破。
原文摘要 · Abstract (English)
The visual world is fundamentally compositional. Visual scenes are defined by the composition of objects and their relations. Hence, it is essential for computer vision systems to reflect and exploit this compositionality to achieve robust and generalizable scene understanding. While major strides have been made toward the development of general-purpose, multimodal generative models, including both text-to-image models and multimodal vision-language models, it remains unclear whether these systems are capable of accurately generating and interpreting scenes involving the composition of multiple objects and relations. In this work, we present an evaluation of the compositional visual processing capabilities in the current generation of text-to-image (DALL-E 3) and multimodal vision-language models (GPT-4V, GPT-4o, Claude Sonnet 3.5, QWEN2-VL-72B, and InternVL2.5-38B), and compare the performance of these systems to human participants. The results suggest that these systems display some ability to solve compositional and relational tasks, showing notable improvements over the previous generation of multimodal models, but with performance nevertheless well below the level of human participants, particularly for more complex scenes involving many ($>5$) objects and multiple relations. These results highlight the need for further progress toward compositional understanding of visual scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。