arXiv:2507.13061cs.CV2025-07中稿 · ACMMM2025被引 1

用分层核心集选择提升视觉语言模型对复杂大场景的理解能力

Advancing Complex Wide-Area Scene Understanding with Hierarchical Coresets Selection

  • 通过理论保障的重要度函数分层筛选关键区域
  • 无需微调即可在任意尺度下快速理解新场景
  • 适配任意视觉语言模型,通用性强

场景理解是计算机视觉的核心任务之一,旨在从图像中提取语义信息以识别物体、场景类别及其相互关系。尽管视觉语言模型(VLMs)取得了进展,但在适应未见的复杂大范围场景方面仍面临挑战。为此,本文提出分层核心集选择(HCS)机制,通过理论保障的重要度函数(考虑效用、代表性、鲁棒性和协同性)逐步优化所选区域,使VLM在不需额外微调的情况下,仅利用最少可解释区域即可实现对任意尺度未见场景的快速理解,并缓解特征密度不足问题。HCS为即插即用方法,兼容任意VLM。实验表明,该方法在多种任务中均展现出优越性能与广泛适用性。

原文摘要 · Abstract (English)

Scene understanding is one of the core tasks in computer vision, aiming to extract semantic information from images to identify objects, scene categories, and their interrelationships. Although advancements in Vision-Language Models (VLMs) have driven progress in this field, existing VLMs still face challenges in adaptation to unseen complex wide-area scenes. To address the challenges, this paper proposes a Hierarchical Coresets Selection (HCS) mechanism to advance the adaptation of VLMs in complex wide-area scene understanding. It progressively refines the selected regions based on the proposed theoretically guaranteed importance function, which considers utility, representativeness, robustness, and synergy. Without requiring additional fine-tuning, HCS enables VLMs to achieve rapid understandings of unseen scenes at any scale using minimal interpretable regions while mitigating insufficient feature density. HCS is a plug-and-play method that is compatible with any VLM. Experiments demonstrate that HCS achieves superior performance and universality in various tasks.

场景理解视觉语言模型核心集选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。