揭示了上下文学习中任务信息剔除的内在机制
Mechanism of Task-oriented Information Removal in In-context Learning
- 通过低秩滤波选择性移除隐藏层冗余信息,引导模型聚焦目标任务
- 少样本上下文学习实质是模拟这种信息剔除过程,提升输出准确率
- 发现关键的去噪注意力头,其失效会显著降低模型性能
上下文学习(ICL)是基于现代语言模型的一种新兴少样本学习范式,但其内部机制仍不清晰。本文从信息移除的新视角展开研究。在零样本场景下,语言模型将查询编码为包含所有可能任务信息的非选择性表示,导致输出随意,准确率接近零。我们发现,通过低秩滤波从隐藏状态中选择性移除特定信息,可有效引导模型聚焦于目标任务。基于此,通过设计度量指标观察隐藏状态,发现少样本ICL实质上模拟了这种任务导向的信息移除过程:从纠缠的非选择性表示中剔除冗余信息,结合示范改进输出,构成ICL的核心机制。此外,我们识别出引发该移除操作的关键注意力头,称为去噪头(Denoising Heads)。通过消融实验阻断该移除过程,模型在正确标签未出现在示范中的情况下,ICL准确率显著下降,验证了信息移除机制及去噪头的关键作用。
原文摘要 · Abstract (English)
In-context Learning (ICL) is an emerging few-shot learning paradigm based on modern Language Models (LMs), yet its inner mechanism remains unclear. In this paper, we investigate the mechanism through a novel perspective of information removal. Specifically, we demonstrate that in the zero-shot scenario, LMs encode queries into non-selective representations in hidden states containing information for all possible tasks, leading to arbitrary outputs without focusing on the intended task, resulting in near-zero accuracy. Meanwhile, we find that selectively removing specific information from hidden states by a low-rank filter effectively steers LMs toward the intended task. Building on these findings, by measuring the hidden states on carefully designed metrics, we observe that few-shot ICL effectively simulates such task-oriented information removal processes, selectively removing the redundant information from entangled non-selective representations, and improving the output based on the demonstrations, which constitutes a key mechanism underlying ICL. Moreover, we identify essential attention heads inducing the removal operation, termed Denoising Heads, which enables the ablation experiments blocking the information removal operation from the inference, where the ICL accuracy significantly degrades, especially when the correct label is absent from the few-shot demonstrations, confirming both the critical role of the information removal mechanism and denoising heads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。