揭示任务向量如何在上下文学习中生成并发挥作用
Understanding Task Vectors in In-Context Learning: Emergence, Functionality, and Limitations
- 任务向量是多个示范通过线性组合形成的单一表示
- 在高秩映射任务中任务向量会失效,实验证实此现象
- 适合研究大模型上下文学习机制的学者参考
任务向量为加速上下文学习(ICL)推理提供了一种高效机制,通过将任务特定信息压缩为单一可复用表示。尽管其在实践中表现优异,但其产生原理与功能机制仍不明确。本文提出线性组合猜想:任务向量本质是原始示范通过线性组合形成的一个单一示范。我们通过损失曲面分析证明,在三元组格式提示下训练的线性变换器中,任务向量自然涌现。进一步预测任务向量无法有效表征高秩映射,并在实际大语言模型上验证了该失败现象。通过显著性分析与参数可视化,我们发现向少样本提示中注入多个任务向量可提升性能。结果深化了对任务向量的理解,揭示了基于变压器模型的ICL内在机制。
原文摘要 · Abstract (English)
Task vectors offer a compelling mechanism for accelerating inference in in-context learning (ICL) by distilling task-specific information into a single, reusable representation. Despite their empirical success, the underlying principles governing their emergence and functionality remain unclear. This work proposes the Linear Combination Conjecture, positing that task vectors act as single in-context demonstrations formed through linear combinations of the original ones. We provide both theoretical and empirical support for this conjecture. First, we show that task vectors naturally emerge in linear transformers trained on triplet-formatted prompts through loss landscape analysis. Next, we predict the failure of task vectors on representing high-rank mappings and confirm this on practical LLMs. Our findings are further validated through saliency analyses and parameter visualization, suggesting an enhancement of task vectors by injecting multiple ones into few-shot prompts. Together, our results advance the understanding of task vectors and shed light on the mechanisms underlying ICL in transformer-based models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。