arXiv:2511.19317cs.CLcs.AI2025-11

构建超5万条多领域孟加拉语摘要数据集,提升文本摘要系统泛化能力

MultiBanAbs: A Comprehensive Multi-Domain Bangla Abstractive Text Summarization Dataset

  • 整合博客、报纸等多源文本,覆盖多样写作风格
  • 包含54,000+文章与摘要对,涵盖新闻、影评等多领域
  • 适配低资源语言研究,推动孟加拉语NLP发展

本研究构建了一个全新的孟加拉语抽象式文本摘要数据集,旨在从多样化来源的孟加拉语文章中生成简洁摘要。现有研究多集中于风格固定的新闻文本,难以适应真实世界中内容形式的多样性。在数字时代,博客、报纸和社交媒体持续产生海量孟加拉语内容,亟需能缓解信息过载的摘要系统。为此,我们收集了超过54,000篇来自Cinegolpo等博客及Samakal、The Business Standard等报纸的文章与对应摘要,构建了跨领域、多风格的数据集。相比单一领域资源,该数据集更具适应性和实用性。我们采用LSTM、BanglaT5-small和MTS-small等深度学习与迁移学习模型进行训练与评估,验证其作为未来孟加拉语自然语言处理研究基准的潜力。该数据集为构建鲁棒摘要系统奠定基础,并助力低资源语言的NLP资源拓展。

原文摘要 · Abstract (English)

This study developed a new Bangla abstractive summarization dataset to generate concise summaries of Bangla articles from diverse sources. Most existing studies in this field have concentrated on news articles, where journalists usually follow a fixed writing style. While such approaches are effective in limited contexts, they often fail to adapt to the varied nature of real-world Bangla texts. In today's digital era, a massive amount of Bangla content is continuously produced across blogs, newspapers, and social media. This creates a pressing need for summarization systems that can reduce information overload and help readers understand content more quickly. To address this challenge, we developed a dataset of over 54,000 Bangla articles and summaries collected from multiple sources, including blogs such as Cinegolpo and newspapers such as Samakal and The Business Standard. Unlike single-domain resources, our dataset spans multiple domains and writing styles. It offers greater adaptability and practical relevance. To establish strong baselines, we trained and evaluated this dataset using several deep learning and transfer learning models, including LSTM, BanglaT5-small, and MTS-small. The results highlight its potential as a benchmark for future research in Bangla natural language processing. This dataset provides a solid foundation for building robust summarization systems and helps expand NLP resources for low-resource languages.

文本摘要孟加拉语多领域低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。