评测大模型在孟加拉语新闻分类中的表现,发现通义千问效果最佳。
Bengali Text Classification: An Evaluation of Large Language Model Approaches
- 用指令微调的大模型直接处理孟加拉语文本分类
- 通义千问2.5达到72%准确率,体育类表现最好
- 适合资源匮乏语言的NLP研究者参考
孟加拉语文本分类是自然语言处理的重要任务,但因缺乏大规模标注数据集和预训练模型而面临挑战。本研究评估了三种指令微调的大语言模型(LLaMA 3.1 8B Instruct、LLaMA 3.2 3B Instruct 和 Qwen 2.5 7B Instruct)在孟加拉语新闻文章分类中的表现。实验基于来自 Kaggle 的数据集,涵盖主要孟加拉国报纸 Prothom Alo 的文章。结果显示,Qwen 2.5 在相同分类框架下取得最高准确率 72%,在“体育”类别中尤为突出;相比之下,LLaMA 3.1 和 LLaMA 3.2 的准确率分别为 53% 和 56%。研究证明大模型在孟加拉语文本分类中具有潜力,即使资源有限。未来工作将探索更多模型、解决类别不平衡问题并优化微调策略以提升性能。
原文摘要 · Abstract (English)
Bengali text classification is a Significant task in natural language processing (NLP), where text is categorized into predefined labels. Unlike English, Bengali faces challenges due to the lack of extensive annotated datasets and pre-trained language models. This study explores the effectiveness of large language models (LLMs) in classifying Bengali newspaper articles. The dataset used, obtained from Kaggle, consists of articles from Prothom Alo, a major Bangladeshi newspaper. Three instruction-tuned LLMs LLaMA 3.1 8B Instruct, LLaMA 3.2 3B Instruct, and Qwen 2.5 7B Instruct were evaluated for this task under the same classification framework. Among the evaluated models, Qwen 2.5 achieved the highest classification accuracy of 72%, showing particular strength in the "Sports" category. In comparison, LLaMA 3.1 and LLaMA 3.2 attained accuracies of 53% and 56%, respectively. The findings highlight the effectiveness of LLMs in Bengali text classification, despite the scarcity of resources for Bengali NLP. Future research will focus on exploring additional models, addressing class imbalance issues, and refining fine-tuning approaches to improve classification performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。