arXiv:2502.05390cs.CLcs.LG2025-02ACL被引 12

用注意力头加权生成任务向量,让大模型跨模态泛化更稳定

Learning Task Representations from In-Context Learning

  • 通过梯度优化注意力头权重,构建可解释的任务向量
  • 在文本与回归任务中均保持任务信息完整性
  • 提出新基准评估跨模态任务泛化能力

大型语言模型在上下文学习(ICL)中表现出色,可通过示例提示适应新任务而无需参数更新。然而,任务如何在模型内部编码与泛化仍不明确。为此,我们提出一种自动化方法,将任务信息编码为变压器架构中注意力头的加权和,权重通过因果梯度下降优化。现有方法在文本以外模态上泛化效果不佳。为此,我们设计了一个基准,用于评估任务向量在函数回归任务中是否能保持任务保真度。所提方法成功从上下文示例中提取特定任务信息,在文本与回归任务中表现优异,展现出跨模态泛化能力。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated remarkable proficiency in in-context learning (ICL), where models adapt to new tasks through example-based prompts without requiring parameter updates. However, understanding how tasks are internally encoded and generalized remains a challenge. To address some of the empirical and technical gaps in the literature, we introduce an automated formulation for encoding task information in ICL prompts as a function of attention heads within the transformer architecture. This approach computes a single task vector as a weighted sum of attention heads, with the weights optimized causally via gradient descent. Our findings show that existing methods fail to generalize effectively to modalities beyond text. In response, we also design a benchmark to evaluate whether a task vector can preserve task fidelity in functional regression tasks. The proposed method successfully extracts task-specific information from in-context demonstrations and excels in both text and regression tasks, demonstrating its generalizability across modalities.

上下文学习任务表征大模型跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。