arXiv:2504.19627cs.CLcs.AI2025-04被引 8

让视觉语言模型像人一样用概念理解图像,大幅降低计算开销。

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning

  • 通过隐式对比学习构建无需标注的视觉概念模型。
  • 在LLaVA-1.5-7B上减少85%计算量,性能基本不变。
  • 适合需要高效视觉理解的机器人、智能系统应用。

大型视觉语言模型(LVLMs)因其强大的视觉-语言推理能力,在具身智能等实际任务中至关重要。然而,现有LVLMs以标记级别处理整张图像,效率远低于人类基于概念的信息分析方式,这种低效源于缺乏视觉概念建模能力,限制了其在真实场景的应用。为此,我们提出VCM,一种端到端自监督的视觉概念建模框架。VCM利用多实例隐式对比学习与视觉-语言指令微调,构建视觉概念模型,无需昂贵的概念级标注。实验表明,VCM显著降低计算成本(如在LLaVA-1.5-7B上减少85%的浮点运算次数),同时在多样化的图像理解任务中保持优异性能。此外,VCM提升了视觉编码器在经典视觉概念感知任务中的表现。大量定量与定性实验验证了VCM的有效性与高效性。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) are pivotal for real-world AI tasks like embodied intelligence due to their strong vision-language reasoning abilities. However, current LVLMs process entire images at the token level, which is inefficient compared to humans who analyze information and generate content at the conceptual level, extracting relevant visual concepts with minimal effort. This inefficiency, stemming from the lack of a visual concept model, limits LVLMs' usability in real-world applications. To address this, we propose VCM, an end-to-end self-supervised visual concept modeling framework. VCM leverages implicit contrastive learning across multiple sampled instances and vision-language fine-tuning to construct a visual concept model without requiring costly concept-level annotations. Our results show that VCM significantly reduces computational costs (e.g., 85\% fewer FLOPs for LLaVA-1.5-7B) while maintaining strong performance across diverse image understanding tasks. Moreover, VCM enhances visual encoders' capabilities in classic visual concept perception tasks. Extensive quantitative and qualitative experiments validate the effectiveness and efficiency of VCM.

视觉概念自监督学习模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。