arXiv:2410.01261cs.CV2024-10中稿 · CVPR被引 4

提升视觉语言模型对被遮挡物体的理解能力

OCC-MLLM:Empowering Multimodal Large Language Model For the Understanding of Occluded Objects

  • 设计新视觉编码器专攻遮挡物体识别
  • 构建大规模含遮挡物体的图文数据集
  • 适合研究遮挡场景下多模态理解的学者

现有大规模视觉语言多模态模型在理解遮挡物体方面存在不足。当前最先进的多模态模型无法通过通用视觉编码器在视觉-语言任务中有效描述遮挡物体。另一个挑战是缺乏包含大量遮挡物体的图像-文本配对数据集。为此,我们提出一种新型多模态模型,采用新设计的视觉编码器以理解RGB图像中的遮挡物体,并构建了一个大规模用于训练和理解遮挡物体的视觉-语言配对数据集。我们通过实验与最先进模型进行对比。

原文摘要 · Abstract (English)

There is a gap in the understanding of occluded objects in existing large-scale visual language multi-modal models. Current state-of-the-art multimodal models fail to provide satisfactory results in describing occluded objects for visual-language multimodal models through universal visual encoders. Another challenge is the limited number of datasets containing image-text pairs with a large number of occluded objects. Therefore, we introduce a novel multimodal model that applies a newly designed visual encoder to understand occluded objects in RGB images. We also introduce a large-scale visual-language pair dataset for training large-scale visual-language multimodal models and understanding occluded objects. We start our experiments comparing with the state-of-the-art models.

多模态遮挡理解视觉编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。