开源视觉大模型因缺乏层级知识,难以识别生物分类关系。
The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
- 构建百万级多选视觉问答任务,基于六类生物分类体系
- 模型在识别鱼类等层级概念时表现差,准确率显著低于预期
- 适合研究视觉-语言模型对知识层次理解的学者参考
本文揭示,许多开源大语言模型(LLMs)缺乏对视觉世界层级结构的认知,甚至不了解基本的生物学分类体系。这一缺陷使它们成为视觉大模型进行层级视觉识别(如识别海葵鱼但不识别脊椎动物)的瓶颈。研究通过构建约一百万条四选一视觉问答(VQA)任务,涵盖六种分类体系和四种图像数据集得出此结论。有趣的是,使用这些VQA任务微调视觉大模型后,发现仍强化了LLM的瓶颈效应——其层级一致性提升幅度远大于视觉模型自身。作者推测,只有当基础语言模型具备相应分类知识,开源视觉大模型才能真正掌握视觉概念的层级关系。
原文摘要 · Abstract (English)
This paper reveals that many open-source large language models (LLMs) lack hierarchical knowledge about our visual world, unaware of even well-established biology taxonomies. This shortcoming makes LLMs a bottleneck for vision LLMs' hierarchical visual recognition (e.g., recognizing Anemone Fish but not Vertebrate). We arrive at these findings using about one million four-choice visual question answering (VQA) tasks constructed from six taxonomies and four image datasets. Interestingly, finetuning a vision LLM using our VQA tasks reaffirms LLMs' bottleneck effect because the VQA tasks improve the LLMs' hierarchical consistency more than the vision LLMs'. We conjecture that one cannot make open-source vision LLMs understand visual concepts hierarchically until LLMs possess corresponding taxonomy knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。