arXiv:2505.23911cs.CL2025-05被引 2

研究发现大模型靠多个任务向量实现少样本学习,单一向量不够用。

One Task Vector is not Enough: A Large-Scale Study for In-Context Learning

  • 用3096个任务验证大模型少样本学习机制
  • 中间层(如第15层)任务向量效果最佳,复杂任务需多个向量
  • 揭示任务知识分布存储,适合研究模型内部机理的学者

上下文学习(ICL)使大语言模型(LLM)能通过少量示例适应新任务,其中任务向量——特定隐藏状态激活——被认为编码了任务信息。现有研究受限于小规模基准,难以全面分析。我们引入新数据集QuiteAFew,包含3,096个多样化的少样本任务,每个任务有30个输入输出对,源自Alpaca数据集。在Llama-3-8B上实验表明:(1) 任务向量性能在中间层(如第15层)达到峰值;(2) 效果随任务类型显著变化;(3) 复杂任务依赖多个子任务特异的向量,而非单一向量,暗示任务知识以分布式方式表征。

原文摘要 · Abstract (English)

In-context learning (ICL) enables Large Language Models (LLMs) to adapt to new tasks using few examples, with task vectors - specific hidden state activations - hypothesized to encode task information. Existing studies are limited by small-scale benchmarks, restricting comprehensive analysis. We introduce QuiteAFew, a novel dataset of 3,096 diverse few-shot tasks, each with 30 input-output pairs derived from the Alpaca dataset. Experiments with Llama-3-8B on QuiteAFew reveal: (1) task vector performance peaks at an intermediate layer (e.g., 15th), (2) effectiveness varies significantly by task type, and (3) complex tasks rely on multiple, subtask-specific vectors rather than a single vector, suggesting distributed task knowledge representation.

上下文学习任务向量大模型机理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。