arXiv:2601.10117cs.CV2026-01

多提示融合与排列优化,提升图像修复模型的快速适应能力

Beyond Single Prompts: Synergistic Fusion and Arrangement for VICL

  • 多提示自适应融合,保留互补信息
  • 不同排列设计轻量MLP,提升结构感知
  • 双向微调增强模块协作,跨任务泛化强

视觉上下文学习(VICL)使图像修复模型仅通过少量提示即可快速适应新视觉任务。但现有方法存在两大问题:(1)仅选择最相似提示,忽略其他高质量提示的互补信息;(2)未能利用不同提示排列所隐含的结构信息。本文提出一个端到端的VICL框架,首先通过自适应融合模块聚合多个提示的关键模式与标注,生成更精确的上下文提示;其次引入特定排列的轻量MLP,将布局先验解耦于主模型,对整体影响极小;此外,采用双向微调机制,交换查询与提示角色,促使模型从融合上下文中重建原始提示,从而增强融合模块与修复模型间的协同。在前景分割、单目标检测和图像着色任务上的实验表明,该方法表现更优,且具备强大的跨任务泛化能力。

原文摘要 · Abstract (English)

Vision In-Context Learning (VICL) enables inpainting models to quickly adapt to new visual tasks from only a few prompts. However, existing methods suffer from two key issues: (1) selecting only the most similar prompt discards complementary cues from other high-quality prompts; and (2) failing to exploit the structured information implied by different prompt arrangements. We propose an end-to-end VICL framework to overcome these limitations. Firstly, an adaptive Fusion Module aggregates critical patterns and annotations from multiple prompts to form more precise contextual prompts. Secondly, we introduce arrangement-specific lightweight MLPs to decouple layout priors from the core model, while minimally affecting the overall model. In addition, an bidirectional fine-tuning mechanism swaps the roles of query and prompt, encouraging the model to reconstruct the original prompt from fused context and thus enhancing collaboration between the fusion module and the inpainting model. Experiments on foreground segmentation, single-object detection, and image colorization demonstrate superior results and strong cross-task generalization of our method.

视觉上下文学习图像修复多提示融合结构化提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。