arXiv:2503.03062cs.CLcs.AI2025-03被引 5

用自生成标注实现上千次上下文学习,性能超越真实标签

Scaling Laws for Many-Shot In-Context Learning with Self-Generated Annotations

  • 用自动生成标注替代真实标签,构建可扩展的上下文学习框架
  • 超过1000个示范时性能达最优,验证了规模化潜力
  • 提出迭代伪标签方法,分类任务提升最高达6.8%

为降低高质量标注数据的成本,研究者提出用自生成标注替代真实标签进行上下文学习(ICL)。尽管已有方法在少样本场景表现良好,但难以扩展至多样本场景。本文基于半监督学习框架,构建包含标注生成、示范选择和上下文推理的流程,提出一种简单基线,在零样本、少样本和多样本设置下均优于使用真实标签的ICL。值得注意的是,该基线存在显著缩放规律:当示范数量超过1000时达到最佳性能。为进一步挖掘半监督ICL的多样本潜力,提出IterPSD方法,融合迭代优化与课程伪标签技术,在分类任务上实现最高6.8%的性能提升。

原文摘要 · Abstract (English)

The high cost of obtaining high-quality annotated data for in-context learning (ICL) has motivated the development of methods that use self-generated annotations in place of ground-truth labels. While these approaches have shown promising results in few-shot settings, they generally do not scale to many-shot scenarios. In this work, we study ICL with self-generated examples using a framework analogous to traditional semi-supervised learning, consisting of annotation generation, demonstration selection, and in-context inference. Within this framework, we propose a simple baseline that outperforms ground-truth ICL in zero-shot, few-shot, and many-shot settings. Notably, we observe a scaling law with this baseline, where optimal performance is achieved with more than 1,000 demonstrations. To fully exploit the many-shot capabilities of semi-supervised ICL, we introduce IterPSD, an iterative annotation approach that integrates iterative refinement and curriculum pseudo-labeling techniques from semi-supervised learning, yielding up to 6.8% additional gains on classification tasks.

上下文学习自生成标注缩放定律半监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。