arXiv:2504.14200cs.CVcs.AI2025-04被引 7

用视觉特征优化图像分类的少样本学习,提升效果且省资源

Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization

  • 以视觉特征为关键点动态构建紧凑演示集
  • 在多个图像分类任务上平均提升超20%性能
  • 适合计算资源受限的实时应用场景

上下文学习(ICL)使大视觉语言模型(LVLMs)无需参数更新即可适应新任务,仅需从大量支持集中选取少量示例。然而,筛选有信息量的示例带来高计算与内存开销。现有方法虽尝试在文本分类中选取小型代表性子集(coreset),但评估所有支持集样本仍耗时,且丢弃样本会造成信息损失。这些方法在图像分类中效果更差,因特征空间差异。为此,本文提出基于关键点的共集优化(KeCO),利用未使用数据构建紧凑且信息丰富的共集。引入视觉特征作为共集中的键,作为不同选择策略下更新样本的锚点。通过利用支持集中的未使用样本更新已选共集样本的键,使随机初始化的共集以低计算成本演化为更优共集。在粗粒度与细粒度图像分类基准上的大量实验表明,KeCO显著提升了图像分类任务的ICL性能,平均提升超过20%。尤其在模拟在线场景下的评估显示其强表现,凸显该框架在资源受限真实场景中的实用价值。

原文摘要 · Abstract (English)

In-context learning (ICL) enables Large Vision-Language Models (LVLMs) to adapt to new tasks without parameter updates, using a few demonstrations from a large support set. However, selecting informative demonstrations leads to high computational and memory costs. While some methods explore selecting a small and representative coreset in the text classification, evaluating all support set samples remains costly, and discarded samples lead to unnecessary information loss. These methods may also be less effective for image classification due to differences in feature spaces. Given these limitations, we propose Key-based Coreset Optimization (KeCO), a novel framework that leverages untapped data to construct a compact and informative coreset. We introduce visual features as keys within the coreset, which serve as the anchor for identifying samples to be updated through different selection strategies. By leveraging untapped samples from the support set, we update the keys of selected coreset samples, enabling the randomly initialized coreset to evolve into a more informative coreset under low computational cost. Through extensive experiments on coarse-grained and fine-grained image classification benchmarks, we demonstrate that KeCO effectively enhances ICL performance for image classification task, achieving an average improvement of more than 20\%. Notably, we evaluate KeCO under a simulated online scenario, and the strong performance in this scenario highlights the practical value of our framework for resource-constrained real-world scenarios.

多模态学习少样本学习图像分类核心集优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。