对比多模态与纯视觉模型在小样本下的表现,发现后者更适应实际小数据场景。
Mind the (Data) Gap: Evaluating Vision Systems in Small Data Applications
- 用NeWT基准测试不同模型在数百到数千样本下的性能
- 纯视觉模型在10个以上样本时持续提升,而多模态模型早期就停滞
- 呼吁研究界重视小数据评估,推动理论与应用对接
实际计算机视觉任务的应用依赖于小数据场景(数百至数千个标注样本),这类场景在生态监测、医学诊断和工业质检等需昂贵专家标注的领域至关重要。然而我们发现,当前计算机视觉研究忽视了这一小数据范式,评估日益集中于零样本和少样本学习。本文利用自然世界任务(NeWT)基准,比较多模态大语言模型(MLLMs)与纯视觉方法在不同训练集规模下的表现。结果显示,MLLMs表现出早期性能平缓,而纯视觉方法在整个小数据范围内持续提升,当训练样本超过10个时性能差距显著扩大。本文首次系统性地对比了两类方法在小数据环境中的表现,并倡导在人工智能研究中明确开展小数据评估,以更好实现理论进展与实际部署的衔接。
原文摘要 · Abstract (English)
The practical application of AI tools for specific computer vision tasks relies on the "small-data regime" of hundreds to thousands of labeled samples. This small-data regime is vital for applications requiring expensive expert annotations, such as ecological monitoring, medical diagnostics or industrial quality control. We find, however, that computer vision research has ignored the small data regime as evaluations increasingly focus on zero- and few-shot learning. We use the Natural World Tasks (NeWT) benchmark to compare multi-modal large language models (MLLMs) and vision-only methods across varying training set sizes. MLLMs exhibit early performance plateaus, while vision-only methods improve throughout the small-data regime, with performance gaps widening beyond 10 training examples. We provide the first comprehensive comparison between these approaches in small-data contexts and advocate for explicit small-data evaluations in AI research to better bridge theoretical advances with practical deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。