首个大规模平衡波斯语社交媒体文本数据集,助力精准分类研究。
PerSoMed: A Large-Scale Balanced Dataset for Persian Social Media Text Classification
- 构建9类共3.6万条平衡数据,覆盖经济到科技等主题。
- 采用混合标注与增强策略,显著提升模型性能,最佳达F1 0.962。
- 适合做波斯语NLP、社会舆情分析及用户行为建模的研究者使用。
本研究首次提出一个大规模、类别均衡的波斯语社交媒体文本分类数据集,旨在填补该领域资源匮乏的空白。数据集包含9个类别(经济、艺术、体育、政治、社会、健康、心理、历史、科技),每类4,000条样本,总计36,000条。原始数据来自多个波斯语社交平台的60,000条帖子,经严格预处理后,结合ChatGPT少样本提示与人工验证完成混合标注。为缓解类别不平衡,采用语义冗余去除的欠采样及融合词法替换与生成式提示的数据增强策略。在多种模型上进行基准测试,包括BiLSTM、XLM-RoBERTa(LoRA与AdaLoRA适配)、FaBERT、SBERT架构及波斯语专用TookaBERT(Base与Large)。实验表明,基于Transformer的模型持续优于传统神经网络,其中TookaBERT-Large表现最佳(精确率:0.9622,召回率:0.9621,F1得分:0.9621)。各类别评估显示整体性能稳健,但社会与政治类文本因固有模糊性得分略低。本研究不仅提供高质量数据集,还对前沿模型进行了全面评估,为波斯语NLP在趋势分析、社会行为建模和用户分类等方面的发展奠定坚实基础。数据集已公开可用。
原文摘要 · Abstract (English)
This research introduces the first large-scale, well-balanced Persian social media text classification dataset, specifically designed to address the lack of comprehensive resources in this domain. The dataset comprises 36,000 posts across nine categories (Economic, Artistic, Sports, Political, Social, Health, Psychological, Historical, and Science & Technology), each containing 4,000 samples to ensure balanced class distribution. Data collection involved 60,000 raw posts from various Persian social media platforms, followed by rigorous preprocessing and hybrid annotation combining ChatGPT-based few-shot prompting with human verification. To mitigate class imbalance, we employed undersampling with semantic redundancy removal and advanced data augmentation strategies integrating lexical replacement and generative prompting. We benchmarked several models, including BiLSTM, XLM-RoBERTa (with LoRA and AdaLoRA adaptations), FaBERT, SBERT-based architectures, and the Persian-specific TookaBERT (Base and Large). Experimental results show that transformer-based models consistently outperform traditional neural networks, with TookaBERT-Large achieving the best performance (Precision: 0.9622, Recall: 0.9621, F1- score: 0.9621). Class-wise evaluation further confirms robust performance across all categories, though social and political texts exhibited slightly lower scores due to inherent ambiguity. This research presents a new high-quality dataset and provides comprehensive evaluations of cutting-edge models, establishing a solid foundation for further developments in Persian NLP, including trend analysis, social behavior modeling, and user classification. The dataset is publicly available to support future research endeavors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。