arXiv:2503.10233cs.CL2025-03被引 1

用Longformer改进波斯语长文档摘要,数据量达30万篇。

ARLED: Leveraging LED-based ARMAN Model for Abstractive Summarization of Persian Long Documents

  • 基于Longformer的ARMAN模型处理长文本
  • 在30万篇波斯语论文上实现良好摘要效果
  • 适合需要处理波斯语文本的研究者

文本数据量持续增长,学者们难以高效阅读和理解长篇研究文献。自动摘要技术可将长文档压缩为简明信息,其中抽象式摘要通过理解文本语义生成更连贯内容。尽管预训练模型如BERT、BART、T5在多种语言中取得进展,长文档摘要仍具挑战。本文聚焦波斯语,构建了包含30万篇全文的波斯语论文数据集,采用基于Longformer架构的ARMAN模型进行抽象式摘要生成。实验结果表明该方法在波斯语文本摘要任务中表现优异,论文还系统回顾了相关工作,详述方法与实验,并展望未来方向。

原文摘要 · Abstract (English)

The increasing volume of textual data poses challenges in reading and comprehending large documents, particularly for scholars who need to extract useful information from research articles. Automatic text summarization has emerged as a powerful tool to condense lengthy documents into concise and informative summaries. Depending on the approach used, text summarization can be categorized as either extractive or abstractive. While extractive methods are commonly used due to their simplicity, they often miss important information. On the other hand, Abstractive Summarization can generate more coherent and informative summaries by understanding the underlying meaning of the text. Abstractive techniques have gained attention in various languages, and recent advancements have been achieved through pre-training models such as BERT, BART, and T5. However, the challenge of summarizing long documents remains, and alternative models like Longformer have been introduced to address this limitation. In this context, this paper focuses on abstractive summarization in the Persian language. The authors introduce a new dataset of 300,000 full-text Persian papers obtained from the Ensani website and apply the ARMAN model, based on the Longformer architecture, to generate summaries. The experimental results demonstrate promising performance in Persian text summarization. The paper provides a comprehensive overview of related work, discusses the methodology, presents the experimental results, and concludes with future research directions.

摘要生成波斯语长文本ARMAN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。