用概率模型解释大模型如何通过上下文示例快速学习。
A Theoretical Interpretation of In-Context Learning via Probabilistic Modeling

- 构建概率模型分析上下文学习机制
- 推导出不同分布下性能的理论表达式
- 揭示示例数量与相似性对效果的影响
上下文学习(ICL)是一种新兴范式,利用大语言模型(LLM)中固有的语义信息来生成用户查询的答案。尽管ICL表现出色,但其通用建模和严格理论分析仍不充分。本文提出一种概率模型用于描述ICL,并推导了在一般参数分布和指数族分布下的性能结果。基于这些推导,文章解释了多个因素对ICL性能的影响,包括示范数量、概率模型对参数变化的敏感度,以及示范与查询之间的相似性。
原文摘要 · Abstract (English)
In-context learning (ICL) is an emerging paradigm that employs the semantic information inherent in large language models (LLMs) for generating answers to user queries. While the remarkable performance of ICL has been widely known, a general modeling and a rigorous theoretical analysis of this paradigm are still lacking. This work presents a probabilistic model for ICL and derives the performance of ICL for both general parametric distributions and exponential families. Based on the derived results, the work explains the impact of multiple factors such as the number of demonstrations, the sensitivity of the probabilistic model to the variation of its parameters, as well as the similarity between the demonstrations and the query on the performance of ICL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。