构建细粒度视觉评估基准,诊断大模型在识别细节时的短板。
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: From Evaluation to Diagnosis

- 设计涵盖百万级问题的多场景细粒度评测集,支持语义与视觉特征双维度评估。
- 发现当前大视觉语言模型在细粒度任务中仍表现不足,失败源于视觉表征、语义对齐等多重瓶颈。
- 提供可诊断模型缺陷的分析框架,适合研究模型鲁棒性与数据设计的学者使用。
近年来,大视觉语言模型(LVLMs)展现出卓越的多模态感知与推理能力。尽管已有诸多基准从整体或特定任务角度评估其性能,但它们在细粒度图像任务——计算机视觉的基础——中的能力仍不明确。为此,我们提出FG-BMK,一个包含101万问题和28万张图像的综合性细粒度评估基准,覆盖从通用物体到专业领域的多样化场景。FG-BMK通过人机双范式,联合评估对话级细粒度语义识别与特征级视觉可区分性,实现对模型失败原因的诊断:是视觉表示不足、视觉-语义对齐弱,还是细粒度知识有限。在多种代表性LVLM/VLM上进行的广泛实验表明,当前模型在细粒度识别上仍存在明显不足,失败由视觉表征、语义对齐、模态对齐及类别知识等多重瓶颈交织导致。我们进一步分析了训练设计因素对提升细粒度能力的影响,并考察了视觉与语言扰动对预测结果的影响。这些发现为理解现有模型局限提供了诊断视角,也为未来数据构建与模型设计提供了指导。代码已开源,地址为https://fg-bmk.github.io/。
原文摘要 · Abstract (English)
Recent advancements in Large Vision-Language Models (LVLMs) have demonstrated remarkable multimodal perception and reasoning capabilities. While numerous benchmarks have evaluated LVLMs from holistic or task-specific perspectives, their capabilities on fine-grained image tasks-fundamental to computer vision-remain insufficiently understood. To address this gap, we introduce FG-BMK, a comprehensive fine-grained evaluation benchmark containing 1.01 million questions and 0.28 million images, covering diverse scenarios from common object-centric domains to specialized domains. FG-BMK jointly evaluates dialogue-level fine-grained semantic recognition and feature-level visual discriminability through human-oriented and machine-oriented paradigms, enabling diagnostic analysis of whether LVLM failures arise from insufficient visual representations, weak visual-to-semantic grounding, or limited fine-grained knowledge. Through extensive experiments on a diverse set of representative LVLMs/VLMs, we find that current LVLMs remain inadequate fine-grained recognizers, with failures arising from intertwined bottlenecks in visual representations, semantic grounding, modality alignment, and category-level knowledge. We further analyze training design factors for improving fine-grained capabilities and examine how visual and linguistic perturbations affect LVLM predictions. These findings provide diagnostic insights into the limitations of current LVLMs and offer guidance for future data construction and model design in developing more reliable LVLMs for fine-grained visual tasks. Our code is open-source and available at https://fg-bmk.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。