研究大模型如何捕捉数千词跨度的长程上下文信息。
How much do contextualized representations encode long-range context?
- 用扰动实验和几何度量分析表示层对长程内容的编码能力。
- 不同模型在相同困惑度下表现差异大,源于长程上下文编码程度不同。
- 混合架构比纯循环模型更有效捕捉全序列结构,适合长文本任务。
我们分析神经自回归语言模型中的上下文表征,重点关注跨越数千个标记的长程上下文。方法采用扰动设置与 extit{各向异性校准余弦相似度}度量,从表示几何角度捕捉长程模式的上下文化程度。以标准解码器仅变压器为例,发现相似困惑度可导致显著不同的下游任务性能,这可归因于长程内容上下文化程度的差异。进一步扩展至其他模型,涵盖新型架构与训练配置,结果表明各类架构对高复杂度(即不可压缩)序列的编码能力下降;全循环模型严重依赖局部上下文,而混合模型更有效地编码整个序列结构。初步分析模型规模与训练配置对长程上下文编码的影响,提示改进现有语言模型的潜在方向。
原文摘要 · Abstract (English)
We analyze contextual representations in neural autoregressive language models, emphasizing long-range contexts that span several thousand tokens. Our methodology employs a perturbation setup and the metric \emph{Anisotropy-Calibrated Cosine Similarity}, to capture the degree of contextualization of long-range patterns from the perspective of representation geometry. We begin the analysis with a case study on standard decoder-only Transformers, demonstrating that similar perplexity can exhibit markedly different downstream task performance, which can be explained by the difference in contextualization of long-range content. Next, we extend the analysis to other models, covering recent novel architectural designs and various training configurations. The representation-level results illustrate a reduced capacity for high-complexity (i.e., less compressible) sequences across architectures, and that fully recurrent models rely heavily on local context, whereas hybrid models more effectively encode the entire sequence structure. Finally, preliminary analysis of model size and training configurations on the encoding of long-range context suggest potential directions for improving existing language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。