arXiv:2604.22884cs.CVcs.AI2026-04

首个针对多模态大模型小物体理解能力的评测基准

Can Multimodal Large Language Models Truly Understand Small Objects?

论文配图:Can Multimodal Large Language Models Truly Understand Small Objects?
图 1 · 摘自论文原文
  • 构建自动化的视觉问答生成策略,创建1.8万组小物体图文数据
  • 15个主流模型在驾驶/航拍/水下场景中均表现不佳
  • 新训练数据提升模型对小物体的理解能力,适合视觉感知研究者

多模态大语言模型在图像视频分析、数理竞赛等任务中展现出潜力,但在小物体理解(SOU)任务上仍空白。为此,我们提出SOUBench,首个全面评估现有MLLM小物体理解能力的基准。首先设计高效自动化视觉问答生成方法,构建包含18,204组VQA对的新数据集,涵盖6个子任务和驾驶、航拍、水下三种典型场景。随后对15个前沿MLLM进行全面评估,揭示其在小物体理解上的薄弱表现。进一步开发包含11,226组VQA对的SOU-Train多模态训练数据集,通过监督微调最新MLLM,验证其显著提升小物体理解能力。实验表明,SOUBench、SOU-VQA与SOU-Train共同为社区提供了关键实证基础,推动具备更强小物体理解能力模型的发展。数据集与代码:https://github.com/Hanfj-X/SOU。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have shown promising potential in diverse understanding tasks, e.g., image and video analysis, math and physics olympiads. However, they remain blank and unexplored for Small Object Understanding (SOU) tasks. To fill this gap, we introduce SOUBench, the first and comprehensive benchmark for exploring the small objects understanding capability of existing MLLMs. Specifically, we first design an effective and automatic visual question-answer generation strategy, constructing a new SOU-VQA evaluation dataset, with 18,204 VQA pairs, six relevant sub-tasks, and three dominant scenarios (i.e., Driving, Aerial, and Underwater). Then, we conduct a comprehensive evaluation on 15 state-of-the-art MLLMs and reveal their weak capabilities in small object understanding. Furthermore, we develop SOU-Train, a multimodal training dataset with 11,226 VQA pairs, to improve the SOU capabilities of MLLMs. Through supervising fine-tuning of the latest MLLM, we demonstrate that SOU-Train can effectively enhance the latest MLLM's ability to understand small objects. Comprehensive experimental results demonstrate that, the proposed SOUBench, along with the SOU-VQA and SOU-Train datasets, provides a crucial empirical foundation to the community for further developing models with enhanced small object understanding capabilities. Datasets and Code: https://github.com/Hanfj-X/SOU.

小物体理解多模态模型评测基准视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。