将上下文学习扩展为涵盖多种语言模型自适应能力的统一框架
The broader spectrum of in-context learning
- 提出上下文学习的新视角:任何能降低后续预测损失的序列分布都算在上下文学习
- 揭示语言模型在指代消解、时间序列外推等任务中的潜在学习机制
- 强调泛化能力应从任务学习、呈现形式灵活度、应用迁移多维度评估
语言模型通过少量示例实现上下文学习的能力引发广泛关注。本文提出一种新视角,将这种监督式少样本学习置于更广泛的元学习型上下文学习谱系中。我们认为,只要上下文能非平凡地降低后续预测的损失,即可视为某种形式的上下文学习。这一视角有助于统一解释语言模型表现出的多种能力,如根据指令或角色扮演调整任务、时间序列外推等。该视角还揭示了上下文学习可能根植于低层语言依赖处理机制(如指代消解或平行结构)。此外,该观点强调泛化的重要性,建议从多个维度研究:不仅包括学习新内容的能力,还包括从不同呈现方式学习及应用所学的能力。本文还讨论了与元学习、目标条件代理等领域的联系,并指出未来研究应考虑上下文学习的更广谱系及其多样化泛化类型。
原文摘要 · Abstract (English)
The ability of language models to learn a task from a few examples in context has generated substantial interest. Here, we provide a perspective that situates this type of supervised few-shot learning within a much broader spectrum of meta-learned in-context learning. Indeed, we suggest that any distribution of sequences in which context non-trivially decreases loss on subsequent predictions can be interpreted as eliciting a kind of in-context learning. We suggest that this perspective helps to unify the broad set of in-context abilities that language models exhibit -- such as adapting to tasks from instructions or role play, or extrapolating time series. This perspective also sheds light on potential roots of in-context learning in lower-level processing of linguistic dependencies (e.g. coreference or parallel structures). Finally, taking this perspective highlights the importance of generalization, which we suggest can be studied along several dimensions: not only the ability to learn something novel, but also flexibility in learning from different presentations, and in applying what is learned. We discuss broader connections to past literature in meta-learning and goal-conditioned agents, and other perspectives on learning and adaptation. We close by suggesting that research on in-context learning should consider this broader spectrum of in-context capabilities and types of generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。