arXiv:2608.20964cs.CLcs.AI2026-08

SAraBERT提升阿拉伯语文档摘要质量,用新评估指标确保核心思想全覆盖。

Extractive Summarization for Arabic Documents Using SAraBERT with a Semantic Siamese Similarity Evaluation Metric

论文配图:Extractive Summarization for Arabic Documents Using SAraBERT with a Semantic Siamese Similarity Evaluation Metric
图 1 · 摘自论文原文
  • 在AraBERT基础上加入句间交互层,增强摘要提取能力。
  • 新提出的语义暹罗相似度指标显著提升摘要与原文的语义覆盖度。
  • 适合关注阿拉伯语信息压缩与多语言自然语言处理的研究者。

本研究提出SAraBERT,是AraBERT的增强版本,专为抽取式摘要任务设计,引入句间Transformer层以捕捉句子间语义关系。为确保生成摘要充分覆盖文档核心思想,我们提出一种新型评估指标——语义暹罗相似度(Semantic Siamese Similarity),用于衡量两段文本间的语义相似性。通过BLEU、ROUGE及该新指标对SAraBERT及同类模型进行验证,仿真结果表明所提模型在摘要质量上表现优异,具有显著有效性,可推动后续相关研究。

原文摘要 · Abstract (English)

In this research, we introduce SAraBERT, an enhanced version of AraBERT which proposes inter-sentence transformer layers for extractive summarization tasks. To ensure that the summaries generated by SAraBERT achieve a high coverage of the document's main ideas, we propose Semantic Siamese Similarity, a novel evaluation metric that measures the level of similarity between two text inputs. We validated using BLEU, ROUGE, and Semantic Siamese similarity on Sarabert and published related models. Simulation results showed the effectiveness of our proposed model and motivate follow on research.

阿拉伯语摘要生成语义评估Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。