arXiv:2501.06256cs.CLcs.AI2025-01被引 4

发现文本外模态实现上下文学习的关键机制

Unlocking In-Context Learning for Natural Datasets Beyond Language Modelling

  • 通过训练数据中的重复标记促进上下文学习
  • 在视觉和脑电数据集上成功激活上下文学习能力
  • 适合研究多模态模型快速适应的新方法

大语言模型具备上下文学习(ICL)能力,可在不更新权重的情况下仅凭上下文示例完成新任务。尽管ICL在自然语言任务中表现良好,但在文本以外的模态中其出现机制尚不明确。本文系统揭示了支持自回归模型及多种模态下ICL涌现的关键属性:训练数据序列中的精确标记重复是关键因素,能提升ICL的稳定性和减少性能波动;同时强调训练任务难度对ICL生成的重要性。基于这些新见解,我们在多个视觉数据集和更具挑战性的脑电分类任务上成功解锁了ICL能力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit In-Context Learning (ICL), which enables the model to perform new tasks conditioning only on the examples provided in the context without updating the model's weights. While ICL offers fast adaptation across natural language tasks and domains, its emergence is less straightforward for modalities beyond text. In this work, we systematically uncover properties present in LLMs that support the emergence of ICL for autoregressive models and various modalities by promoting the learning of the needed mechanisms for ICL. We identify exact token repetitions in the training data sequences as an important factor for ICL. Such repetitions further improve stability and reduce transiency in ICL performance. Moreover, we emphasise the significance of training task difficulty for the emergence of ICL. Finally, by applying our novel insights on ICL emergence, we unlock ICL capabilities for various visual datasets and a more challenging EEG classification task.

上下文学习多模态模型泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。