arXiv:2504.14630cs.CL2025-04

首个索拉尼库尔德语科研文献自动摘要研究,填补语言资源空白

Automatic Text Summarization (ATS) for Research Documents in Sorani Kurdish

  • 基于231篇库尔德语科研论文构建数据集与模型
  • 引入句权重与TF-IDF算法,最高准确率达19.58%
  • 为库尔德语NLP研究提供可复用的资源与方法

从科学文献中提取简洁信息有助于学习者、研究人员和实践者。自动文本摘要(ATS)是自然语言处理(NLP)的重要应用,可自动化该过程。尽管多种语言已有相关方法,但库尔德语因资源匮乏仍处于发展初期。本研究基于伊拉克库尔德地区两所大学四个院系的231篇索拉尼库尔德语科研论文,平均每篇26页,构建了数据集与语言模型。采用句权重与词频-逆文档频率(TF-IDF)算法进行两次实验,分别包含与不包含结论部分,平均词数分别为5,492.3与5,266.96。结果通过人工与自动评估(ROUGE-1、ROUGE-2、ROUGE-L)验证,最佳准确率为19.58%。六位专家按三项标准进行人工评估,结果因文档而异。本研究为库尔德语NLP领域的自动摘要及相关方向提供了宝贵资源。

原文摘要 · Abstract (English)

Extracting concise information from scientific documents aids learners, researchers, and practitioners. Automatic Text Summarization (ATS), a key Natural Language Processing (NLP) application, automates this process. While ATS methods exist for many languages, Kurdish remains underdeveloped due to limited resources. This study develops a dataset and language model based on 231 scientific papers in Sorani Kurdish, collected from four academic departments in two universities in the Kurdistan Region of Iraq (KRI), averaging 26 pages per document. Using Sentence Weighting and Term Frequency-Inverse Document Frequency (TF-IDF) algorithms, two experiments were conducted, differing in whether the conclusions were included. The average word count was 5,492.3 in the first experiment and 5,266.96 in the second. Results were evaluated manually and automatically using ROUGE-1, ROUGE-2, and ROUGE-L metrics, with the best accuracy reaching 19.58%. Six experts conducted manual evaluations using three criteria, with results varying by document. This research provides valuable resources for Kurdish NLP researchers to advance ATS and related fields.

自动摘要库尔德语NLP数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。