arXiv:2511.08977cs.CV2025-11中稿 · AAAI被引 2

用核心集方法高效选示例,提升视觉语言模型的上下文学习效果。

Efficient and Effective In-context Demonstration Selection with Coreset

  • 基于核心集构建多样示例子集,提升信息效用。
  • 在多个数据集上显著优于随机、相似度等传统方法。
  • 适合需要高效准确示例选择的视觉语言模型应用。

上下文学习(ICL)已成为大型视觉语言模型(LVLMs)的强大范式,使模型能直接利用输入上下文中的少量示例进行推理。然而,该方法的有效性高度依赖于示例的选择,而这一过程是NP难问题。传统策略如随机采样、基于相似度的采样和infoscore采样常导致效率低下或性能不佳,难以兼顾效率与效果。本文提出一种名为基于核心集的双重检索(CoDR)的新框架。我们证明,具有多样性的子集中样本可实现更高的期望互信息。为此,引入聚类剪枝方法构建与查询更匹配且保持多样性的核心集。同时,设计双重检索机制,在保证效率的同时实现全局示例选择。实验表明,该方法显著优于现有策略,为有效且高效的示例选择提供了稳健解决方案。

原文摘要 · Abstract (English)

In-context learning (ICL) has emerged as a powerful paradigm for Large Visual Language Models (LVLMs), enabling them to leverage a few examples directly from input contexts. However, the effectiveness of this approach is heavily reliant on the selection of demonstrations, a process that is NP-hard. Traditional strategies, including random, similarity-based sampling and infoscore-based sampling, often lead to inefficiencies or suboptimal performance, struggling to balance both efficiency and effectiveness in demonstration selection. In this paper, we propose a novel demonstration selection framework named Coreset-based Dual Retrieval (CoDR). We show that samples within a diverse subset achieve a higher expected mutual information. To implement this, we introduce a cluster-pruning method to construct a diverse coreset that aligns more effectively with the query while maintaining diversity. Additionally, we develop a dual retrieval mechanism that enhances the selection process by achieving global demonstration selection while preserving efficiency. Experimental results demonstrate that our method significantly improves the ICL performance compared to the existing strategies, providing a robust solution for effective and efficient demonstration selection.

上下文学习核心集视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。