arXiv:2410.01861cs.CV2024-10中稿 · ECCV被引 2

让大模型看懂被遮挡物体,自监督学习提升描述准确率16.92%。

OCC-MLLM-Alpha:Empowering Multi-modal Large Language Model for the Understanding of Occluded Objects with Self-Supervised Test-Time Learning

  • 基于3D生成与自监督测试时学习,增强对遮挡物体的理解能力。
  • 在SOMVideo数据集上比现有模型提升16.92%的描述准确率。
  • 适合需要精准理解复杂场景中遮挡目标的研究者与开发者。

现有大规模视觉语言多模态模型在理解遮挡物体方面存在不足。当前最先进的多模态模型无法通过通用视觉编码器和监督学习策略有效描述遮挡物体。为此,我们提出一种多模态大语言模型框架及配套的自监督学习策略,并引入3D生成支持。我们在大规模数据集SOMVideo [18] 上进行实验对比,初始结果表明,该方法相比最先进视觉语言模型(VLM)提升16.92%。该研究为提升模型在遮挡场景下的语义理解能力提供了新路径。

原文摘要 · Abstract (English)

There is a gap in the understanding of occluded objects in existing large-scale visual language multi-modal models. Current state-of-the-art multi-modal models fail to provide satisfactory results in describing occluded objects through universal visual encoders and supervised learning strategies. Therefore, we introduce a multi-modal large language framework and corresponding self-supervised learning strategy with support of 3D generation. We start our experiments comparing with the state-of-the-art models in the evaluation of a large-scale dataset SOMVideo [18]. The initial results demonstrate the improvement of 16.92% in comparison with the state-of-the-art VLM models.

多模态遮挡理解自监督学习3D生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。