构建首个细粒度视觉语言评估基准,揭示大模型在细节识别上的短板。
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation
- 设计包含百万级问题的细粒度图文评测集FG-BMK。
- 发现训练方式与模态对齐显著影响模型细粒度识别能力。
- 适合研究多模态模型局限性与细粒度视觉任务的研究者。
近年来,大型视觉语言模型(LVLMs)在多模态感知方面展现出卓越能力,受到广泛关注。尽管已有大量评估研究,但针对细粒度图像任务——计算机视觉的基础任务——的系统性评估仍严重不足。为填补这一空白,我们提出了一个全面的细粒度评估基准FG-BMK,包含101万条问题和33万张图像。我们的评估从人本与机器双视角出发,聚焦模型在语义识别与细粒度特征表征方面的能力。通过对十二个代表性LVLM/VLM的广泛实验,我们揭示了训练范式、模态对齐、扰动敏感性及细粒度类别推理对性能的关键影响。本工作为当前LVLM的局限性提供了重要洞察,并为未来数据构建与模型设计提供了指导。代码已开源,详见https://github.com/SEU-VIPGroup/FG-BMK。
原文摘要 · Abstract (English)
Recent advancements in Large Vision-Language Models (LVLMs) have demonstrated remarkable multimodal perception capabilities, garnering significant attention. While numerous evaluation studies have emerged, assessing LVLMs both holistically and on specialized tasks, fine-grained image tasks-fundamental to computer vision-remain largely unexplored. To fill this gap, we introduce a comprehensive fine-grained evaluation benchmark, i.e., FG-BMK, comprising 1.01 million questions and 0.33 million images. Our evaluation systematically examines LVLMs from both human-oriented and machine-oriented perspectives, focusing on their semantic recognition and fine-grained feature representation capabilities. Through extensive experiments on twelve representative LVLMs/VLMs, we uncover key findings regarding the influence of training paradigms, modality alignment, perturbation susceptibility, and fine-grained category reasoning on task performance. This work provides critical insights into the limitations of current LVLMs and offers guidance for future data construction and model design in the development of more advanced LVLMs. Our code is open-source and available at https://github.com/SEU-VIPGroup/FG-BMK.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。