arXiv:2511.18618cs.CLcs.AI2025-11被引 3

用混合模型同时分析孟加拉语新闻标题分类与情感,提升信息理解效率。

A Unified BERT-CNN-BiLSTM Framework for Simultaneous Headline Classification and Sentiment Analysis of Bangla News

  • 融合BERT、CNN与BiLSTM的混合模型,提升文本理解能力。
  • 在9014条新闻上实现标题分类78.57%、情感分析73.43%准确率。
  • 首个同时处理孟加拉语新闻分类与情感分析的研究,适合低资源语言任务。

日常生活中,报纸是影响公众讨论的重要信息来源。然而,从不同报纸和在线新闻平台中高效获取新闻内容仍具挑战。新闻标题的情感分析能揭示新闻主题(如政治、体育)及情绪倾向(正面、负面、中性),帮助快速把握新闻基调。本研究提出一种基于自然语言处理的先进方法,结合使用BERT-CNN-BiLSTM混合迁移学习模型,对孟加拉语新闻标题进行分类与情感分析。我们采用名为BAN-ABSA的9014条新闻标题数据集,首次实现孟加拉语新闻标题与情感的联合分类。针对数据不平衡问题,实验采用两种策略:策略1在分割前进行过采样与欠采样,获得最高性能,标题分类78.57%,情感分析73.43%;策略2在原始不平衡数据上直接训练,标题分类达81.37%,情感分析64.46%。所提模型显著优于所有基线模型,在孟加拉语新闻分类与情感分析任务上达到新基准,验证了联合利用标题与情感数据的重要性,为低资源语言文本分类提供有力支持。

原文摘要 · Abstract (English)

In our daily lives, newspapers are an essential information source that impacts how the public talks about present-day issues. However, effectively navigating the vast amount of news content from different newspapers and online news portals can be challenging. Newspaper headlines with sentiment analysis tell us what the news is about (e.g., politics, sports) and how the news makes us feel (positive, negative, neutral). This helps us quickly understand the emotional tone of the news. This research presents a state-of-the-art approach to Bangla news headline classification combined with sentiment analysis applying Natural Language Processing (NLP) techniques, particularly the hybrid transfer learning model BERT-CNN-BiLSTM. We have explored a dataset called BAN-ABSA of 9014 news headlines, which is the first time that has been experimented with simultaneously in the headline and sentiment categorization in Bengali newspapers. Over this imbalanced dataset, we applied two experimental strategies: technique-1, where undersampling and oversampling are applied before splitting, and technique-2, where undersampling and oversampling are applied after splitting on the In technique-1 oversampling provided the strongest performance, both headline and sentiment, that is 78.57\% and 73.43\% respectively, while technique-2 delivered the highest result when trained directly on the original imbalanced dataset, both headline and sentiment, that is 81.37\% and 64.46\% respectively. The proposed model BERT-CNN-BiLSTM significantly outperforms all baseline models in classification tasks, and achieves new state-of-the-art results for Bangla news headline classification and sentiment analysis. These results demonstrate the importance of leveraging both the headline and sentiment datasets, and provide a strong baseline for Bangla text classification in low-resource.

文本分类情感分析低资源语言BERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。