构建首个大规模马拉地语摘要数据集并训练专用BART模型
L3Cube-MahaSum: A Comprehensive Dataset and BART Models for Abstractive Text Summarization in Marathi
- 用25000篇新闻构建马拉地语摘要数据集,人工验证摘要质量
- 基于MahaSUM训练的IndicBART在抽象摘要任务中表现优异
- 开源数据集与模型,助力印地语系语言NLP研究
我们提出MahaSUM数据集,这是一个大规模、多样化的马拉地语新闻文章集合,用于促进印地语系语言中抽取式摘要任务的模型训练与评估。该数据集包含25,000个样本,通过从多个在线新闻源抓取并人工验证摘要而构建。此外,我们使用MahaSUM数据集训练了一个专为印地语系语言设计的IndicBART模型。我们在抽象摘要任务上评估了所训练模型的性能,证明其能生成高质量的马拉地语摘要。本工作推动了印地语系语言自然语言处理研究的发展,并为未来基于先进模型的研究提供了宝贵资源。数据集和模型已公开发布于https://github.com/l3cube-pune/MarathiNLP。
原文摘要 · Abstract (English)
We present the MahaSUM dataset, a large-scale collection of diverse news articles in Marathi, designed to facilitate the training and evaluation of models for abstractive summarization tasks in Indic languages. The dataset, containing 25k samples, was created by scraping articles from a wide range of online news sources and manually verifying the abstract summaries. Additionally, we train an IndicBART model, a variant of the BART model tailored for Indic languages, using the MahaSUM dataset. We evaluate the performance of our trained models on the task of abstractive summarization and demonstrate their effectiveness in producing high-quality summaries in Marathi. Our work contributes to the advancement of natural language processing research in Indic languages and provides a valuable resource for future research in this area using state-of-the-art models. The dataset and models are shared publicly at https://github.com/l3cube-pune/MarathiNLP
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。