arXiv:2509.21926cs.CV2025-09被引 2

解决视觉上下文学习中依赖单一示例的偏见问题

PANICL: Mitigating Over-Reliance on Single Prompt in Visual In-Context Learning

  • 用多示例邻居机制替代单示例引导,平滑预测得分
  • 在多种任务上超越强基线,提升稳定性和鲁棒性
  • 无需训练即可适配多种模型,适合实际部署

视觉上下文学习(VICL)通过输入输出图像对(即上下文对)作为提示,指导模型完成多样视觉任务。然而,当前方法常过度依赖单一上下文对,导致预测偏差与不稳定。本文提出无需训练的PAtch-based k-Nearest neighbor VICL(PANICL),通过利用多个上下文对来缓解该问题。PANICL在不同上下文对间平滑分配分数,降低偏倚,无需额外训练。大量实验表明,在前景分割、单目标检测、着色、多目标分割和关键点检测等任务上均优于强基线。此外,PANICL在数据集级迁移(如从COCO到Pascal)和标签空间迁移(如FSS-1000)下仍保持强鲁棒性,并可通用适配SegGPT、Painter、LVM等其他VICL模型,展现良好泛化能力与广泛适用性。

原文摘要 · Abstract (English)

Visual In-Context Learning (VICL) uses input-output image pairs, referred to as in-context pairs (or examples), as prompts alongside query images to guide models in performing diverse vision tasks. However, VICL often suffers from over-reliance on a single in-context pair, which can lead to biased and unstable predictions. We introduce PAtch-based $k$-Nearest neighbor visual In-Context Learning (PANICL), a general training-free framework that mitigates this issue by leveraging multiple in-context pairs. PANICL smooths assignment scores across pairs, reducing bias without requiring additional training. Extensive experiments on a variety of tasks, including foreground segmentation, single object detection, colorization, multi-object segmentation, and keypoint detection, demonstrate consistent improvements over strong baselines. Moreover, PANICL exhibits strong robustness to domain shifts, including dataset-level shift (e.g., from COCO to Pascal) and label-space shift (e.g., FSS-1000), and generalizes well to other VICL models such as SegGPT, Painter, and LVM, highlighting its versatility and broad applicability.

视觉上下文学习多示例学习模型鲁棒性无训练框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。