arXiv:2508.00385cs.CLcs.LG2025-08被引 1

提出梯度流选例法,让有效示范更突出,提升大模型上下文学习效果。

Multi-Layer Attention is the Amplifier of Demonstration Effectiveness

  • 用梯度流分析示范有效性,发现模型已学或无关的示范无效。
  • 多层模型会放大示范差异,越往后越聚焦有效示范。
  • 新方法GradS选示范更准,平均比顶尖基线提升6.8%。

众多研究探讨了上下文学习(ICL)有效性的机制,但多数假设示范均有效,而实证表明并非所有示范都能带来性能提升。本文基于梯度流与线性自注意力模型分析示范无效的原因:当示范信息已被模型学习或与用户查询无关时,其失效。进一步发现,在多层模型中,示范间有效性差异随层数增加而被放大,导致模型更关注有效示范。现有示范选择方法多关注与查询的相关性,忽略模型已吸收的信息。为此,本文提出基于梯度流的选例方法GradS,以示范对用户查询的梯度流幅度为选择标准,确保所选示范有效。在四个主流大模型和五个基准数据集上的实验验证了该现象,并证明GradS相较最强基线平均提升6.8%,证实其有效性。

原文摘要 · Abstract (English)

Numerous studies have investigated the underlying mechanisms of in-context learning (ICL) effectiveness to inspire the design of related methods. However, existing work predominantly assumes the effectiveness of the demonstrations provided within ICL, while many research indicates that not all demonstrations are effective, failing to yielding any performance improvement during ICL. Therefore, in this paper, we investigate the reasons behind demonstration ineffectiveness. Our analysis is based on gradient flow and linear self-attention models. By setting the gradient flow to zero, we deduce that a demonstration becomes ineffective if its information has either been learned by the model or is irrelevant to the user query. Furthermore, we demonstrate that in multi-layer models, the disparity in effectiveness among demonstrations is amplified with layer increasing, causing the model to focus more on effective ones. Considering that current demonstration selection methods primarily focus on the relevance to the user query while overlooking the information that the model has already assimilated, we propose a novel method called GradS, which leverages gradient flow for demonstration selection. We use the magnitude of the gradient flow of the demonstration with respect to a given user query as the criterion, thereby ensuring the effectiveness of the chosen ones. We validate our derivation and GradS on four prominent LLMs across five mainstream datasets. The experimental results confirm that the disparity in effectiveness among demonstrations is magnified as the model layer increases, substantiating our derivations. Moreover, GradS achieves a relative improvement of $6.8\%$ on average over the strongest baselines, demonstrating its effectiveness.

大模型上下文学习示范选择梯度流

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。