arXiv:2410.05903cs.CLcs.AI2024-10被引 1

让大模型轻松处理超长文档,突破输入长度限制。

Automatic Summarization of Long Documents

  • 提出三种新算法,无需修改模型结构即可扩展输入长度
  • 在7万词以上文本上,BERTScore显著提升,ROUGE分数竞争力强
  • 适合需要处理长篇文档的科研、法律、新闻等场景

互联网每天新增海量文本数据,其利用与解读日益困难,自动文本摘要因此至关重要,可高效提取关键信息并节省阅读时间。尽管许多基于Transformer的模型在摘要任务中表现优异,但受限于输入长度,无法处理超过其上下文窗口的长文本。本研究提出三种新算法,使任意大语言模型都能在不进行架构修改的情况下,高效突破输入长度限制,充分发挥模型潜力。我们在超过7万词的文本上测试了这些算法,实验结果表明,在保持竞争性ROUGE得分的同时,BERTScore实现显著提升。

原文摘要 · Abstract (English)

A vast amount of textual data is added to the internet daily, making utilization and interpretation of such data difficult and cumbersome. As a result, automatic text summarization is crucial for extracting relevant information, saving precious reading time. Although many transformer-based models excel in summarization, they are constrained by their input size, preventing them from processing texts longer than their context size. This study introduces three novel algorithms that allow any LLM to efficiently overcome its input size limitation, effectively utilizing its full potential without any architectural modifications. We test our algorithms on texts with more than 70,000 words, and our experiments show a significant increase in BERTScore with competitive ROUGE scores.

文本摘要长文档大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。