用大模型反馈直接优化演示选择,提升少样本学习效果
Learning to Select In-Context Demonstration Preferred by Large Language Model
- 通过生成式偏好学习,让LLM直接评价演示好坏
- 19个数据集上优于现有方法,显著提升ICL性能
- 适合需要高效少样本适配的AI研发人员
上下文学习(ICL)使大语言模型在推理时仅用少量示例即可适应新任务。然而,ICL表现高度依赖示例选择。现有检索方法常使用代理目标(如度量学习)进行优化,无法直接提升ICL效果,难以发现真正有效的示例。且当候选池中高质量示例不足时,其判别式检索策略失效。为此,我们提出GenICL,一种利用大语言模型反馈进行生成式偏好学习的新框架,直接优化演示选择以提升ICL表现。在11个任务类别、19个数据集上的实验表明,GenICL在挑选最优演示方面优于现有方法,从而带来更好的ICL性能。
原文摘要 · Abstract (English)
In-context learning (ICL) enables large language models (LLMs) to adapt to new tasks during inference using only a few demonstrations. However, ICL performance is highly dependent on the selection of these demonstrations. Recent work explores retrieval-based methods for selecting query-specific demonstrations, but these approaches often rely on surrogate objectives such as metric learning, failing to directly optimize ICL performance. Consequently, they struggle to identify truly beneficial demonstrations. Moreover, their discriminative retrieval paradigm is ineffective when the candidate pool lacks sufficient high-quality demonstrations. To address these challenges, we propose GenICL, a novel generative preference learning framework that leverages LLM feedback to directly optimize demonstration selection for ICL. Experiments on 19 datasets across 11 task categories demonstrate that GenICL achieves superior performance than existing methods in selecting the most effective demonstrations, leading to better ICL performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。