arXiv:2501.14276cs.CVcs.AI2025-01被引 3

让大模型更懂高分辨率图像,自动聚焦重要区域

Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models

  • 根据语义相关性动态分配子图权重,模拟人眼注意力
  • 在相同参数量下超越同类模型,接近更大模型表现
  • 适合需要高效处理高清图像的多模态应用

随着对大视觉语言模型(LVLMs)高分辨率图像处理需求的增长,子图像划分成为缓解固定分辨率处理导致视觉信息丢失的常用方法。然而,现有划分方法对子图像进行均匀处理,导致图像理解效果不佳。本文发现,与整体图像语义相关性更高的子图像蕴含更丰富的视觉信息,有助于保持模型的视觉理解能力。为此,提出全局语义引导加权模块(GSWA),根据子图像的信息密度动态分配权重,模拟人类视觉注意力机制。该方法使模型聚焦于更具信息量的区域,克服了统一处理的局限性。将GSWA集成至InternVL2-2B框架,构建轻量级高性能模型SleighVL。大量实验表明,SleighVL在参数量相当的情况下优于同类模型,且与更大模型性能相当。本工作为高效、上下文感知的高分辨率图像处理提供了新方向,推动多模态系统发展。

原文摘要 · Abstract (English)

As the demand for high-resolution image processing in Large Vision-Language Models (LVLMs) grows, sub-image partitioning has become a popular approach for mitigating visual information loss associated with fixed-resolution processing. However, existing partitioning methods uniformly process sub-images, resulting in suboptimal image understanding. In this work, we reveal that the sub-images with higher semantic relevance to the entire image encapsulate richer visual information for preserving the model's visual understanding ability. Therefore, we propose the Global Semantic-guided Weight Allocator (GSWA) module, which dynamically allocates weights to sub-images based on their relative information density, emulating human visual attention mechanisms. This approach enables the model to focus on more informative regions, overcoming the limitations of uniform treatment. We integrate GSWA into the InternVL2-2B framework to create SleighVL, a lightweight yet high-performing model. Extensive experiments demonstrate that SleighVL outperforms models with comparable parameters and remains competitive with larger models. Our work provides a promising direction for more efficient and contextually aware high-resolution image processing in LVLMs, advancing multimodal system development.

视觉语言模型高分辨率图像注意力机制子图像分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。