构建尼泊尔语-英语与泰卢固语-英语混合文本数据集,提升低资源语言暴力内容识别能力。
Creating and Evaluating Code-Mixed Nepali-English and Telugu-English Datasets for Abusive Language Detection Using Traditional and Deep Learning Models
- 人工标注2000条泰卢固-英语和5000条尼泊尔-英语混合评论,分两类
- 在多模型上测试,发现深度学习模型在混合语境中表现更优
- 为低资源语言暴力检测提供首个基准数据集,适合社交媒体治理研究
随着多语言用户在社交媒体上的增多,代码混合文本中的暴力语言检测日益困难。代码混合指用户在英文与母语间无缝切换,使攻击性内容常依赖语境或被语言融合掩盖。尽管英语和印地语的暴力检测已有广泛研究,但泰卢固语和尼泊尔语等低资源语言仍缺乏足够数据。本研究构建了包含2000条泰卢固-英语和5000条尼泊尔-英语混合评论的手动标注数据集,分为暴力与非暴力两类,数据源自多个社交平台。经严格预处理后,使用多种机器学习(ML)、深度学习(DL)及大语言模型(LLMs)进行评估,包括逻辑回归、随机森林、支持向量机(SVM)、神经网络(NN)、LSTM、CNN及LLMs,并通过超参数调优、10折交叉验证与显著性检验(t-test)分析性能。研究揭示了代码混合环境下暴力检测的挑战,对比了不同计算方法的效果。该工作为低资源语言的代码混合暴力检测提供了基准,有助于构建更鲁棒的多语言社交媒体治理策略。
原文摘要 · Abstract (English)
With the growing presence of multilingual users on social media, detecting abusive language in code-mixed text has become increasingly challenging. Code-mixed communication, where users seamlessly switch between English and their native languages, poses difficulties for traditional abuse detection models, as offensive content may be context-dependent or obscured by linguistic blending. While abusive language detection has been extensively explored for high-resource languages like English and Hindi, low-resource languages such as Telugu and Nepali remain underrepresented, leaving gaps in effective moderation. In this study, we introduce a novel, manually annotated dataset of 2 thousand Telugu-English and 5 Nepali-English code-mixed comments, categorized as abusive and non-abusive, collected from various social media platforms. The dataset undergoes rigorous preprocessing before being evaluated across multiple Machine Learning (ML), Deep Learning (DL), and Large Language Models (LLMs). We experimented with models including Logistic Regression, Random Forest, Support Vector Machines (SVM), Neural Networks (NN), LSTM, CNN, and LLMs, optimizing their performance through hyperparameter tuning, and evaluate it using 10-fold cross-validation and statistical significance testing (t-test). Our findings provide key insights into the challenges of detecting abusive language in code-mixed settings and offer a comparative analysis of computational approaches. This study contributes to advancing NLP for low-resource languages by establishing benchmarks for abusive language detection in Telugu-English and Nepali-English code-mixed text. The dataset and insights can aid in the development of more robust moderation strategies for multilingual social media environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。