arXiv:2603.18891cs.CVcs.LG2026-03中稿 · ICLR被引 3

提出PromptHub框架,提升多提示视觉上下文学习的融合效果

PromptHub: Enhancing Multi-Prompt Visual In-Context Learning with Locality-Aware Fusion, Concentration and Alignment

  • 基于局部感知融合,利用空间先验捕捉更丰富上下文
  • 通过互补目标联合训练,显著提升三类基础视觉任务性能
  • 适用于分布外场景与多种检索任务,具有强泛化能力

视觉上下文学习(VICL)旨在通过模仿像素级示范完成视觉任务。近期工作提出了提示融合方法,结合多种示范优势,展现出拓展VICL的潜力。然而,现有的基于块的融合框架和模型无关监督限制了有效线索的利用,制约性能提升。为此,我们提出PromptHub框架,通过局部感知融合、集中与对齐机制全面增强多提示学习。PromptHub利用空间先验捕捉更丰富的上下文信息,采用互补的集中、对齐和预测目标相互引导训练,并引入数据增强以强化监督。在三个基础视觉任务上的大量实验验证了PromptHub的优越性。此外,我们还验证了其在分布外设置和多种检索场景下的通用性、可迁移性与鲁棒性。本工作建立了一个可靠的局部感知提示融合范式,超越了以往基于块的方法。代码已公开于 https://github.com/luotc-why/ICLR26-PromptHub。

原文摘要 · Abstract (English)

Visual In-Context Learning (VICL) aims to complete vision tasks by imitating pixel demonstrations. Recent work pioneered prompt fusion that combines the advantages of various demonstrations, which shows a promising way to extend VICL. Unfortunately, the patch-wise fusion framework and model-agnostic supervision hinder the exploitation of informative cues, thereby limiting performance gains. To overcome this deficiency, we introduce PromptHub, a framework that holistically strengthens multi-prompting through locality-aware fusion, concentration and alignment. PromptHub exploits spatial priors to capture richer contextual information, employs complementary concentration, alignment, and prediction objectives to mutually guide training, and incorporates data augmentation to further reinforce supervision. Extensive experiments on three fundamental vision tasks demonstrate the superiority of PromptHub. Moreover, we validate its universality, transferability, and robustness across out-of-distribution settings, and various retrieval scenarios. This work establishes a reliable locality-aware paradigm for prompt fusion, moving beyond prior patch-wise approaches. Code is available at https://github.com/luotc-why/ICLR26-PromptHub.

视觉上下文学习提示融合局部感知多提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。