arXiv:2606.10905cs.CV2026-06

用100万参数小模型挑战视觉上下文学习,发现现有评估体系存在漏洞。

Beyond Model Size: Probing the Gaps in Visual in-Context Learning by Training a Tiny Model

论文配图:Beyond Model Size: Probing the Gaps in Visual in-Context Learning by Training a Tiny Model
图 1 · 摘自论文原文
  • 仅用100万参数和7万张图训练极简视觉模型
  • 小模型在新任务上表现逼近大模型,暴露评估缺陷
  • 适合关注评测标准合理性与模型适应性研究者

视觉上下文学习(VICL)旨在构建可基于少量样本在测试时自适应新任务的视觉模型。当前方法多依赖大规模模型与数据,但其有效性尚未验证。本文训练一个仅有100万参数、使用7万张图像的极小视觉模型,与7000倍更大的基准模型在三种适应场景下对比:(1)分布偏移较小的图像数据,(2)未见过的任务编码,(3)完全新的任务。尽管训练资源差距巨大,小模型在部分任务上表现接近大模型。结果揭示了当前VICL评估体系在任务编码方式、预训练任务选择及评价指标上的系统性盲区,凸显改进适应能力评测方法的迫切需求。

原文摘要 · Abstract (English)

Visual in-Context Learning (VICL) aims at making progress towards adaptive vision models, that can -- based on a few examples -- adapt to a new task at test-time. With the history of in-context learning in natural language processing research, where large, parameter-heavy models are in use, one pathway that current VICL methods take is model- and data-scaling as key ingredients. Yet, it is not clear, whether these ingredients are the key for in-context learning to take shape in vision models. To stress-test such large models, we challenge them with an extreme counterexample: we train a tiny visual in-context model with merely $1$ million parameters and a modest amount of $70,000$ images. We compare the results of this severely capacity capped tiny model to $7,000\times$ larger VICL models in different adaptive settings, (1) on image data with small distribution shifts, (2) on unseen task encodings and (3) on a completely new task, i.e., the setting VICL envisions. With the chasm of training resources between the tiny- and large models, our experiments showcase a lack in how adaptive capabilities are measured, with respect to how tasks are encoded, which tasks were used in pre-training and the choice of metrics. These gaps in current VICL benchmarking underscore a need for innovation in evaluation of adaptive capabilities.

视觉学习模型评估小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。