arXiv:2503.13299cs.CLcs.AI2025-03综述被引 16

梳理Transformer长上下文处理的挑战与方法,助力大模型应对长文本任务。

A Survey on Transformer Context Extension: Approaches and Evaluation

  • 按位置编码、上下文压缩等四类归纳长上下文技术
  • 系统整理现有评估数据集、任务与指标体系
  • 适合关注大模型长文本能力的研究者参考

基于Transformer的大语言模型在自然语言处理中广泛应用,尤其在短文本任务中表现优异。然而,在长上下文场景下,其性能因若干挑战而下降。为此,近期涌现出多种改进方法。本文首先列出预训练大模型处理长上下文的主要挑战,系统回顾相关技术,并提出四类分类体系:位置编码、上下文压缩、检索增强和注意力模式。此外,聚焦长上下文评估,整合现有基准中的数据、任务与评价指标。最后,总结该领域未解问题,并展望未来发展方向。

原文摘要 · Abstract (English)

Large language models (LLMs) based on Transformer have been widely applied in the filed of natural language processing (NLP), demonstrating strong performance, particularly in handling short text tasks. However, when it comes to long context scenarios, the performance of LLMs degrades due to some challenges. To alleviate this phenomenon, there is a number of work proposed recently. In this survey, we first list the challenges of applying pre-trained LLMs to process long contexts. Then systematically review the approaches related to long context and propose our taxonomy categorizing them into four main types: positional encoding, context compression, retrieval augmented, and attention pattern. In addition to the approaches, we focus on the evaluation of long context, organizing relevant data, tasks, and metrics based on existing long context benchmarks. Finally, we summarize unresolved issues in the long context domain and put forward our views on future developments.

Transformer长文本大模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。