解析示例如何影响大模型分类决策,揭示预训练知识与示例的权衡机制。
Towards the Effect of Examples on In-Context Learning: A Theoretical Case Study
- 构建概率模型量化预训练知识、标签频率和噪声对分类的影响。
- 示例数量决定模型更依赖预训练知识还是示例本身。
- 少数类准确率更低,噪声影响取决于具体噪声水平。
上下文学习(ICL)使大语言模型通过少量示例快速适应下游任务,但其机制尚不明确。本文针对二分类任务开展理论研究,提出一种基于高斯混合模型的概率模型,精确量化预训练知识、标签频率和标签噪声对预测准确率的影响。分析表明:当预训练知识与示例知识冲突时,模型是更依赖预训练知识还是示例,取决于示例数量;此外,示例的标签频率和噪声均影响准确性——少数类预测准确率更低,且噪声的影响由两类的具体噪声水平决定。大量仿真验证了理论结果,真实数据实验也与之吻合。本工作揭示了预训练知识与示例在ICL中的作用机制,深化了对大模型分类行为的理解。
原文摘要 · Abstract (English)
In-context learning (ICL) has emerged as a powerful capability for large language models (LLMs) to adapt to downstream tasks by leveraging a few (demonstration) examples. Despite its effectiveness, the mechanism behind ICL remains underexplored. To better understand how ICL integrates the examples with the knowledge learned by the LLM during pre-training (i.e., pre-training knowledge) and how the examples impact ICL, this paper conducts a theoretical study in binary classification tasks. In particular, we introduce a probabilistic model extending from the Gaussian mixture model to exactly quantify the impact of pre-training knowledge, label frequency, and label noise on the prediction accuracy. Based on our analysis, when the pre-training knowledge contradicts the knowledge in the examples, whether ICL prediction relies more on the pre-training knowledge or the examples depends on the number of examples. In addition, the label frequency and label noise of the examples both affect the accuracy of the ICL prediction, where the minor class has a lower accuracy, and how the label noise impacts the accuracy is determined by the specific noise level of the two classes. Extensive simulations are conducted to verify the correctness of the theoretical results, and real-data experiments also align with the theoretical insights. Our work reveals the role of pre-training knowledge and examples in ICL, offering a deeper understanding of LLMs' behaviors in classification tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。