arXiv:2505.06840cs.CV2025-05被引 3

通过聚焦图像关键区域,提升大模型视觉理解效率。

Visual Instruction Tuning with Chain of Region-of-Interest

  • 基于人眼视觉机制,识别并优先处理图像重要区域。
  • 在11个基准上超越LLaVA-NeXT,34B模型超Gemini Pro 1.0六项任务。
  • 适合追求高效高分辨率视觉理解的模型开发者。

高分辨率图像对提升多模态大语言模型(MLLMs)的识别与理解能力至关重要,但直接提高图像分辨率会显著增加计算负担。本文提出链式兴趣区域(CoRoI)方法,通过借鉴人类视觉系统的选择性,识别并优先处理高分辨率图像中的关键信息区域,从而在不处理冗长高分辨率图像标记的情况下,增强多模态视觉理解能力。在11个基准上的广泛实验表明,该方法在7B至34B参数规模下均有效,模型在多种多模态任务中表现优异。尤其值得注意的是,该方法在几乎所有基准上优于LLaVA-NeXT,其34B微调模型在六个基准上超越了专有模型Gemini Pro 1.0,且在MMB、SEED-I和MME任务上优于GPT-4V。

原文摘要 · Abstract (English)

High-resolution (HR) images are pivotal for enhancing the recognition and understanding capabilities of multimodal large language models (MLLMs). However, directly increasing image resolution can significantly escalate computational demands. In this study, we propose a method called Chain of Region-of-Interest (CoRoI) for Visual Instruction Tuning, aimed at alleviating the computational burden associated with high-resolution images for MLLMs. Drawing inspiration from the selective nature of the human visual system, we recognize that not all regions within high-resolution images carry equal importance. CoRoI seeks to identify and prioritize the most informative regions, thereby enhancing multimodal visual comprehension and recognition while circumventing the need for processing lengthy HR image tokens. Through extensive experiments on 11 benchmarks, we validate the efficacy of CoRoI across varying sizes, ranging from 7B to 34B in parameters. Our models consistently demonstrate superior performance across diverse multimodal benchmarks and tasks. Notably, our method outperforms LLaVA-NeXT on almost all benchmarks and our finetuned 34B model surpasses proprietary methods like Gemini Pro 1.0 on six benchmarks, as well as outperforming GPT-4V on MMB, SEED-I, and MME.

视觉理解多模态高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。