对比十年视觉语言模型在复杂社交场景中的表现,发现大模型显著提升描述准确率。
Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models

- 构建复杂社交行为数据集CSB,评估模型对人类互动的描述能力
- 大模型在复杂场景中准确率接近顶尖人类,但空间依赖误差仍存
- 检测、识别与幻觉是影响描述准确性的关键错误类型
过去十年间,视觉语言模型(VLMs)在视觉推理方面取得了显著进展。现有评估多基于简单场景(MS-COCO),缺乏对复杂人类互动与行为的考察,且仅使用少量非筛选的人类描述作为基准,未关注模型错误类型。为此,本文引入复杂社交行为(CSB)数据集,包含100张描绘复杂社会互动的图像。我们分析了2017至2025年间四款预大模型(pre-MLLMs)与五款多模态大语言模型(MLLMs)在场景描述上的演进。在CSB和MS-COCO样本上,评估模型与20位人类描述相对于黄金标准的准确性,并分析五类视觉认知错误:物体检测、识别、幻觉、场景理解与空间依赖。结果表明,相比MS-COCO,CSB数据集显示更显著的性能提升;预大模型准确率远低于最低人类水平,而大模型则达到与顶级人类相当的水平。大模型基本消除了简单与复杂场景间的准确率差距,除偶尔在图像区域选择上与人类不一致(空间依赖误差)外,其余错误类型已大幅减少。检测、识别与幻觉错误对描述准确率影响最大。研究为视觉语言模型十年发展提供了更全面的评估视角。
原文摘要 · Abstract (English)
Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-COCO) that do not showcase complex human interactions or behaviors, only a handful of non-curated human descriptions as a benchmark, and have not focused on understanding the model's error types. Here, we introduce the Complex Social Behavior (CSB) dataset, containing 100 images depicting complex social interactions/behaviors. We analyze the progression of scene descriptions over a decade (2017-2025) of VLMs (four pre-Multimodal Large Language Models, MLLMs, and five MLLMs). We evaluate the accuracy of the models and 20 human descriptions relative to a gold standard on the CSB dataset and on a sample from MS-COCO. We analyzed five visual-cognitive error types: object detection, recognition, hallucination, scene understanding, and spatial dependence. The CSB dataset showed a more pronounced improvement than MS-COCO in scene description accuracy, with pre-MLLMs achieving much lower accuracy than the bottom-ranked human descriptions and MLLMs attaining accuracies similar to the top-ranked human descriptions. We show that MLLMs have eliminated the gap in scene description accuracy between simpler MS-COCO scenes and scenes depicting complex behaviors (CSB). MLLMs have almost eliminated all error types in our tested datasets, except for occasionally relying on different image regions for scene descriptions than humans do (spatial dependence error). We also show that detection, recognition, and hallucination errors have the highest impact on scene description accuracy. Together, our findings provide a more thorough evaluation of how visual language models have advanced over the last decade.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。