arXiv:2601.20129cs.CLcs.AI2026-01被引 1

构建了13.9万条孟加拉语情感数据集,用于仇恨与非仇恨言论分类。

BengaliSent140: A Large-Scale Bengali Binary Sentiment Dataset for Hate and Non-Hate Speech Classification

  • 整合7个现有数据集,统一为二分类标注体系
  • 含139,792条文本,仇恨言论68,548条,非仇恨71,244条
  • 覆盖多领域,适合训练深度学习模型

近年来,孟加拉语情感分析研究日益增多,但受限于大规模多样化的标注数据集稀缺。尽管已有多个公开的孟加拉语情感和仇恨言论数据集,但多数规模较小或局限于单一领域(如社交媒体评论),难以满足现代深度学习模型对海量异构数据的需求。本文提出 BengaliSent140,一个通过整合七个现有孟加拉语文本数据集构建的大规模二分类情感数据集。为保证来源一致性,不同标注体系被系统地统一为二分类格式:非仇恨(0)和仇恨(1)。最终数据集包含139,792条唯一文本样本,其中仇恨言论68,548条,非仇恨言论71,244条,类别分布相对均衡。该数据集融合多源多领域数据,覆盖范围更广,可为深度学习模型的训练与评估提供坚实基础。同时报告了基线实验结果,验证其可用性。数据集已公开发布于 Kaggle。

原文摘要 · Abstract (English)

Sentiment analysis for the Bengali language has attracted increasing research interest in recent years. However, progress remains constrained by the scarcity of large-scale and diverse annotated datasets. Although several Bengali sentiment and hate speech datasets are publicly available, most are limited in size or confined to a single domain, such as social media comments. Consequently, these resources are often insufficient for training modern deep learning based models, which require large volumes of heterogeneous data to learn robust and generalizable representations. In this work, we introduce BengaliSent140, a large-scale Bengali binary sentiment dataset constructed by consolidating seven existing Bengali text datasets into a unified corpus. To ensure consistency across sources, heterogeneous annotation schemes are systematically harmonized into a binary sentiment formulation with two classes: Not Hate (0) and Hate (1). The resulting dataset comprises 139,792 unique text samples, including 68,548 hate and 71,244 not-hate instances, yielding a relatively balanced class distribution. By integrating data from multiple sources and domains, BengaliSent140 offers broader linguistic and contextual coverage than existing Bengali sentiment datasets and provides a strong foundation for training and benchmarking deep learning models. Baseline experimental results are also reported to demonstrate the practical usability of the dataset. The dataset is publicly available at https://www.kaggle.com/datasets/akifislam/bengalisent140/

情感分析仇恨言论孟加拉语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。