用BERT和逻辑回归检测泰米尔语和马拉雅拉姆语社交媒体中的女性针对性辱骂文本
GS_DravidianLangTech@2025: Women Targeted Abusive Texts Detection on Social Media
- 基于BERT与逻辑回归模型,针对南印度语言构建女性歧视文本识别系统
- 在泰米尔语和马拉雅拉姆语数据集上分别达到0.729和0.6279的宏F1分数
- 为低资源语言的性别暴力内容检测提供可复用的技术方案,适合数字安全研究者
社交媒体滥用问题日益严重,亟需有效内容治理技术。本文聚焦于识别针对女性的网络辱骂文本,此类言论旨在伤害或煽动对弱势群体的仇恨。研究采用逻辑回归与BERT作为基础模型,利用DravidianLangTech@2025提供的泰米尔语和马拉雅拉姆语数据集进行训练与评估。实验结果表明,BERT在泰米尔语和马拉雅拉姆语测试集上分别取得0.729和0.6279的宏F1分数,逻辑回归模型表现分别为0.6279和0.5821。该工作为低资源南印度语言中的性别化仇恨言论识别提供了可行方法。
原文摘要 · Abstract (English)
The increasing misuse of social media has become a concern; however, technological solutions are being developed to moderate its content effectively. This paper focuses on detecting abusive texts targeting women on social media platforms. Abusive speech refers to communication intended to harm or incite hatred against vulnerable individuals or groups. Specifically, this study aims to identify abusive language directed toward women. To achieve this, we utilized logistic regression and BERT as base models to train datasets sourced from DravidianLangTech@2025 for Tamil and Malayalam languages. The models were evaluated on test datasets, resulting in a 0.729 macro F1 score for BERT and 0.6279 for logistic regression in Tamil and Malayalam, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。