不同提示方式生成的模型任务表征不同,但可互补提升性能。
Do different prompting methods yield a common task representation in language models?
- 用函数向量分析指令与示范提示下的任务表征差异
- 指令提示也能提取有效函数向量,提升零样本准确率
- 两种提示依赖不同模型组件,适合混合使用
演示和指令是语言模型进行上下文学习(ICL)的两种主要提示方式。相同任务通过不同方式触发,是否产生相似的任务表征?我们通过函数向量(FVs)研究该问题,该方法可提取少样本ICL任务的表征。我们将FVs扩展至短文本指令提示,成功提取出能提升零样本任务准确率的指令函数向量。结果表明,基于演示和基于指令的函数向量利用模型中不同的组件,并通过多个控制实验分离其对性能的贡献。研究发现,不同提示形式并未通过FVs诱导出共同的任务表征,而是引发部分重叠但不同的机制。这为结合指令与演示的实践提供了理论支持,也揭示了跨提示形式统一监控任务推理的挑战,并呼吁进一步探究大模型的任务推理机制。
原文摘要 · Abstract (English)
Demonstrations and instructions are two primary approaches for prompting language models to perform in-context learning (ICL) tasks. Do identical tasks elicited in different ways result in similar representations of the task? An improved understanding of task representation mechanisms would offer interpretability insights and may aid in steering models. We study this through \textit{function vectors} (FVs), recently proposed as a mechanism to extract few-shot ICL task representations. We generalize FVs to alternative task presentations, focusing on short textual instruction prompts, and successfully extract instruction function vectors that promote zero-shot task accuracy. We find evidence that demonstration- and instruction-based function vectors leverage different model components, and offer several controls to dissociate their contributions to task performance. Our results suggest that different task promptings forms do not induce a common task representation through FVs but elicit different, partly overlapping mechanisms. Our findings offer principled support to the practice of combining instructions and task demonstrations, imply challenges in universally monitoring task inference across presentation forms, and encourage further examinations of LLM task inference mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。