提升视觉语言模型对被遮挡物体的理解能力
OCC-MLLM:Empowering Multimodal Large Language Model For the Understanding of Occluded Objects
- 设计新视觉编码器专攻遮挡物体识别
- 构建大规模含遮挡物体的图文数据集
- 适合研究遮挡场景下多模态理解的学者
现有大规模视觉语言多模态模型在理解遮挡物体方面存在不足。当前最先进的多模态模型无法通过通用视觉编码器在视觉-语言任务中有效描述遮挡物体。另一个挑战是缺乏包含大量遮挡物体的图像-文本配对数据集。为此,我们提出一种新型多模态模型,采用新设计的视觉编码器以理解RGB图像中的遮挡物体,并构建了一个大规模用于训练和理解遮挡物体的视觉-语言配对数据集。我们通过实验与最先进模型进行对比。
原文摘要 · Abstract (English)
There is a gap in the understanding of occluded objects in existing large-scale visual language multi-modal models. Current state-of-the-art multimodal models fail to provide satisfactory results in describing occluded objects for visual-language multimodal models through universal visual encoders. Another challenge is the limited number of datasets containing image-text pairs with a large number of occluded objects. Therefore, we introduce a novel multimodal model that applies a newly designed visual encoder to understand occluded objects in RGB images. We also introduce a large-scale visual-language pair dataset for training large-scale visual-language multimodal models and understanding occluded objects. We start our experiments comparing with the state-of-the-art models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。