通过多组合协作融合提升视觉上下文学习的性能
Enhancing Visual In-Context Learning by Multi-Faceted Fusion
- 提出多组合协作融合机制,生成三路互补上下文表征
- 在分割、检测、着色任务上均超越现有方法,提升泛化能力
- 适合需要强泛化与多源信息利用的视觉推理场景
视觉上下文学习(VICL)作为一种新兴范式,使模型能够通过上下文示例完成新视觉任务。主流的‘检索-提示’方法通常仅选取最优单个视觉提示,常忽略其他合适候选者的有价值信息。尽管近期工作尝试将前K个提示融合为单一增强表征,仍将其简化为单一信号,限制了模型推理能力。本文认为需采用更全面、协作式的多面融合以释放多样上下文的潜力。为此,提出新框架:不再将多个提示压缩为单一表征,而是生成三路上下文表征分支,每条由不同组合的高质量提示整合而成。这些互补引导信号输入所提出的MULTI-VQGAN架构,该架构旨在联合解析并利用多源协作信息。在前景分割、单对象检测和图像着色等多样化任务上的大量实验表明,该方法具备强大的跨任务泛化能力、有效的上下文融合效果,并能生成比现有方法更鲁棒、更准确的预测结果。
原文摘要 · Abstract (English)
Visual In-Context Learning (VICL) has emerged as a powerful paradigm, enabling models to perform novel visual tasks by learning from in-context examples. The dominant "retrieve-then-prompt" approach typically relies on selecting the single best visual prompt, a practice that often discards valuable contextual information from other suitable candidates. While recent work has explored fusing the top-K prompts into a single, enhanced representation, this still simply collapses multiple rich signals into one, limiting the model's reasoning capability. We argue that a more multi-faceted, collaborative fusion is required to unlock the full potential of these diverse contexts. To address this limitation, we introduce a novel framework that moves beyond single-prompt fusion towards an multi-combination collaborative fusion. Instead of collapsing multiple prompts into one, our method generates three contextual representation branches, each formed by integrating information from different combinations of top-quality prompts. These complementary guidance signals are then fed into proposed MULTI-VQGAN architecture, which is designed to jointly interpret and utilize collaborative information from multiple sources. Extensive experiments on diverse tasks, including foreground segmentation, single-object detection, and image colorization, highlight its strong cross-task generalization, effective contextual fusion, and ability to produce more robust and accurate predictions than existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。