arXiv:2606.10967cs.CV2026-06

构建跨领域任务的视觉上下文学习基准,揭示现有模型真实适应能力

Quo Vadis, Visual In-Context Learning? A Unified Benchmark Across Domains and Tasks

论文配图:Quo Vadis, Visual In-Context Learning? A Unified Benchmark Across Domains and Tasks
图 1 · 摘自论文原文
  • 设计统一评估框架VIBE,覆盖14个数据集与12种任务
  • 在106种组合上测试6个模型,发现普遍适应性不足
  • 开源工具包,推动更真实、可复现的视觉学习评估

视觉上下文学习被视作实现动态模型的路径,使模型能基于给定上下文生成预测,并在测试时适应新视觉任务。然而,当前对这些模型适应能力的评估局限于与预训练任务或图像领域相似的狭窄场景,实际无需真正适应。为此,我们构建了涵盖多样成像领域和广泛任务的大型视觉上下文学习基准(VIBE)。通过该基准,我们得以更清晰地考察视觉上下文模型面对新图像分布和任务分布时的真实适应能力。我们在14个数据集、12个任务(共106种数据集-任务组合)上对六个模型进行压力测试,采用统一可复现的评估协议,在单样本设置下进行比较。评估揭示了当前视觉上下文学习的关键局限、系统性失败模式及未来有潜力的方向。为促进更广泛的评估,我们将公开发布VIBE工具包。

原文摘要 · Abstract (English)

Visual in-context learning has been proposed as a pathway towards dynamic models that can generate predictions based on a provided context and thereby can adapt to new vision tasks at test-time. Yet, the evaluation of the adaptation capabilities of these models has been limited to narrow setups that mainly mirror tasks or image domains from pre-training for which real adaptation is not required. We address this gap by constructing a broad Visual In-Context BEnchmark (VIBE) with a focus on diverse imaging domains and a wide range of tasks. With this, we are able to get a much clearer picture of the adaptive capabilities of visual in-context models when faced with new image- and task distributions. We stress test six models on $14$ datasets and $12$ tasks (in total, we explore $106$ dataset-task combinations) and compare them under a unified, reproducible evaluation protocol, in an one-shot setting. Our evaluation uncovers key insights on the state of visual in-context learning, including limitations, systematic failure modes and promising directions. To foster broader evaluation, we will openly release our VIBE toolkit.

视觉学习上下文学习基准测试模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。