arXiv:2501.01069cs.CLcs.LG2025-01被引 5

为孟加拉宗教新闻生成新数据集,融合上下文特征提升标题质量

BeliN: A Novel Corpus for Bengali Religious News Headline Generation using Contextual Feature Fusion

  • 提出多输入特征融合模型MultiGen,结合内容与情感、类别等上下文信息
  • 在新数据集BeliN上实现BLEU 18.61、ROUGE-L 24.19,优于仅用内容的基线
  • 面向低资源语言处理研究者,尤其关注宗教新闻与文化语境的中文读者

自动文本摘要,尤其是标题生成,在孟加拉宗教新闻领域仍属关键但研究不足的方向。现有方法多仅依赖文章内容,忽略情感、类别、方面等关键上下文特征,严重制约其效果。本文提出新数据集BeliN(来自孟加拉国主流在线报纸的宗教新闻),并设计MultiGen——一种基于上下文特征融合的多输入标题生成方法。该方法利用BanglaT5、mBART、mT5和mT0等预训练模型,将类别、方面、情感等额外上下文特征与新闻内容融合,有效捕捉传统方法忽视的信息。实验表明,MultiGen在基准方法(仅使用内容)基础上显著提升性能:BLEU得分达18.61(基线16.08),ROUGE-L得分为24.19(基线23.08)。结果凸显了在低资源语言中引入上下文特征的重要性。研究通过弥合语言与文化差距,推动孟加拉语及其他未充分代表语言的自然语言处理发展。为促进可复现性与进一步探索,数据集与代码已公开于https://github.com/akabircs/BeliN。

原文摘要 · Abstract (English)

Automatic text summarization, particularly headline generation, remains a critical yet underexplored area for Bengali religious news. Existing approaches to headline generation typically rely solely on the article content, overlooking crucial contextual features such as sentiment, category, and aspect. This limitation significantly hinders their effectiveness and overall performance. This study addresses this limitation by introducing a novel corpus, BeliN (Bengali Religious News) - comprising religious news articles from prominent Bangladeshi online newspapers, and MultiGen - a contextual multi-input feature fusion headline generation approach. Leveraging transformer-based pre-trained language models such as BanglaT5, mBART, mT5, and mT0, MultiGen integrates additional contextual features - including category, aspect, and sentiment - with the news content. This fusion enables the model to capture critical contextual information often overlooked by traditional methods. Experimental results demonstrate the superiority of MultiGen over the baseline approach that uses only news content, achieving a BLEU score of 18.61 and ROUGE-L score of 24.19, compared to baseline approach scores of 16.08 and 23.08, respectively. These findings underscore the importance of incorporating contextual features in headline generation for low-resource languages. By bridging linguistic and cultural gaps, this research advances natural language processing for Bengali and other underrepresented languages. To promote reproducibility and further exploration, the dataset and implementation code are publicly accessible at https://github.com/akabircs/BeliN.

标题生成低资源语言上下文融合孟加拉语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。