arXiv:2501.09240cs.LG2025-01被引 22

研究提示向量如何在上下文学习中形成并提升模型表现

Task Vectors in In-Context Learning: Emergence, Formation, and Benefit

  • 通过从头训练模型,揭示任务向量的自然生成机制
  • 引入任务向量提示损失,使任务信息强编码于指定位置
  • 提升模型鲁棒性与泛化能力,适合需要可控适配的研究者

上下文学习是变换器模型的一项显著能力,指其能根据简短上下文或历史信息适应特定任务。先前研究发现任务相关信息在模型中局部编码,但因预训练过程不透明,其涌现机制与功能仍不明确。本文在可控环境下,使用从头训练的合成数据集研究任务向量的形成。结果表明,任务向量在特定条件下会自然涌现,但任务信息可能较弱且非局部编码。为促进强任务向量在模型中指定位置的编码,我们提出基于任务向量提示损失(TVP-loss)的辅助训练机制。该方法无需在训练后搜索任务相关编码,显著提升模型的鲁棒性与泛化性能。

原文摘要 · Abstract (English)

In-context learning is a remarkable capability of transformers, referring to their ability to adapt to specific tasks based on a short history or context. Previous research has found that task-specific information is locally encoded within models, though their emergence and functionality remain unclear due to opaque pre-training processes. In this work, we investigate the formation of task vectors in a controlled setting, using models trained from scratch on synthetic datasets. Our findings confirm that task vectors naturally emerge under certain conditions, but the tasks may be relatively weakly and/or non-locally encoded within the model. To promote strong task vectors encoded at a prescribed location within the model, we propose an auxiliary training mechanism based on a task vector prompting loss (TVP-loss). This method eliminates the need to search for task-correlated encodings within the trained model and demonstrably improves robustness and generalization.

上下文学习任务向量模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。