arXiv:2410.08971cs.CL2024-10被引 1

通过关键词触发全局注意力,提升长文档摘要效果

Extra Global Attention Designation Using Keyword Detection in Sparse Transformer Architectures

  • 用关键词前缀激活全局注意力,增强长距离依赖建模
  • 在零样本、少样本和微调场景下均提升摘要质量
  • 适合需要处理长文本的摘要与信息提取任务

本文提出对流行的稀疏Transformer架构Longformer Encoder-Decoder的扩展。稀疏Transformer常难以捕捉文档中长距离上下文关系,如开头与结尾话题的关联。我们提出一种选择性增强全局注意力的方法:在文本前添加关键词,并对这些关键词启用全局注意力。该方法在多个基准数据集上的抽取式摘要任务中均取得提升,涵盖零样本、少样本及微调三种场景。

原文摘要 · Abstract (English)

In this paper, we propose an extension to Longformer Encoder-Decoder, a popular sparse transformer architecture. One common challenge with sparse transformers is that they can struggle with encoding of long range context, such as connections between topics discussed at a beginning and end of a document. A method to selectively increase global attention is proposed and demonstrated for abstractive summarization tasks on several benchmark data sets. By prefixing the transcript with additional keywords and encoding global attention on these keywords, improvement in zero-shot, few-shot, and fine-tuned cases is demonstrated for some benchmark data sets.

稀疏注意力长文本摘要关键词触发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。