arXiv:2412.03565cs.CV2024-12NeurIPS被引 7

通过显式视觉提示微调,提升大模型对图像视频中具体对象的理解能力。

INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning

  • 用显式视觉提示进行指令微调,引导模型关注具体实例。
  • 在多个实例理解基准上性能显著提升,通用图文理解也同步增强。
  • 适合需要精准识别物体的场景,如智能监控、医疗影像分析。

大型多模态模型(LMMs)随着指令微调的发展取得了显著进展。然而,现有模型虽能整体理解图像和视频,但在需要更细粒度认知与对齐的实例级理解上仍表现不足。实例级理解至关重要,因其聚焦于我们最关心的具体元素。令人振奋的是,现有研究发现,当提供显式视觉线索时,最先进的LMMs展现出强大的实例理解能力。受此启发,我们提出Inst-IT,一种通过显式视觉提示指令微调来增强LMMs实例理解能力的方法。Inst-IT包含一个用于诊断多模态实例级理解的基准、一个大规模指令微调数据集,以及一种连续指令微调训练范式,以有效提升现有LMMs在时空实例理解方面的能力。实验结果表明,经Inst-IT增强的模型不仅在Inst-IT基准及其他实例理解基准上表现优异,还在多种通用图像和视频理解基准上实现显著提升。这表明该方法不仅能强化实例级理解,还能全面提升通用图像与视频理解能力。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) have made significant breakthroughs with the advancement of instruction tuning. However, while existing models can understand images and videos at a holistic level, they still struggle with instance-level understanding that requires a more fine-grained comprehension and alignment. Instance-level understanding is crucial for LMMs, as it focuses on the specific elements that we are most interested in. Excitingly, existing works find that the SOTA LMMs exhibit strong instance understanding capabilities when provided with explicit visual cues. Motivated by this, we proposed Inst-IT, a solution to enhance LMMs in Instance understanding via explicit visual prompt Instruction Tuning for instance guidance. Inst-IT consists of a benchmark to diagnose multimodal instance-level understanding, a large-scale instruction-tuning dataset, and a continuous instruction-tuning training paradigm to effectively enhance spatial-temporal instance understanding capabilities of existing LMMs. Experimental results show that, enhanced by Inst-IT, our models not only achieve outstanding performance on Inst-IT Bench and other instance understanding benchmarks, but also demonstrate significant improvements across various generic image and video understanding benchmarks. This highlights that our method not only boosts instance-level understanding but also strengthens the overall capabilities of generic image and video comprehension.

多模态实例理解指令微调视觉提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。