arXiv:2506.16066cs.CL2025-06被引 4

用MURIL模型检测混合印地语英文文本中的网络欺凌,准确率超94%。

Cyberbullying Detection in Hinglish Text Using MURIL and Explainable AI

  • 基于MURIL架构处理印地语英文混写文本的欺凌检测。
  • 在6个数据集上准确率达75%-94%,优于RoBERTa等模型。
  • 支持可解释性分析,适合多语言内容安全研究者使用。

数字通信平台的普及导致全球网络欺凌事件增多,亟需自动化检测系统保护用户。印度语与英语混合(Hinglish)的兴起给现有以单语文本设计的检测系统带来挑战。本文提出基于多语言印度语表征模型MURIL的欺凌检测框架,解决现有方法局限。在六个基准数据集(Bohra et al., BullyExplain, BullySentemo, Kumar et al., HASOC 2021, Mendeley Indo-HateSpeech)上的评估显示,该方法优于RoBERTa和IndicBERT等多语言模型,准确率提升1.36至13.07个百分点,具体为:Bohra数据集86.97%,BullyExplain 84.62%,BullySentemo 86.03%,Kumar数据集75.41%,HASOC 2021 83.92%,Mendeley数据集94.63%。框架包含归因分析与跨语言模式识别的可解释性功能。消融实验表明,选择性层冻结、合理分类头设计及针对混写内容的预处理能提升性能;失败分析揭示了上下文依赖理解、文化认知与跨语言讽刺识别等挑战,为未来多语言欺凌检测研究指明方向。

原文摘要 · Abstract (English)

The growth of digital communication platforms has led to increased cyberbullying incidents worldwide, creating a need for automated detection systems to protect users. The rise of code-mixed Hindi-English (Hinglish) communication on digital platforms poses challenges for existing cyberbullying detection systems, which were designed primarily for monolingual text. This paper presents a framework for cyberbullying detection in Hinglish text using the Multilingual Representations for Indian Languages (MURIL) architecture to address limitations in current approaches. Evaluation across six benchmark datasets -- Bohra \textit{et al.}, BullyExplain, BullySentemo, Kumar \textit{et al.}, HASOC 2021, and Mendeley Indo-HateSpeech -- shows that the MURIL-based approach outperforms existing multilingual models including RoBERTa and IndicBERT, with improvements of 1.36 to 13.07 percentage points and accuracies of 86.97\% on Bohra, 84.62\% on BullyExplain, 86.03\% on BullySentemo, 75.41\% on Kumar datasets, 83.92\% on HASOC 2021, and 94.63\% on Mendeley dataset. The framework includes explainability features through attribution analysis and cross-linguistic pattern recognition. Ablation studies show that selective layer freezing, appropriate classification head design, and specialized preprocessing for code-mixed content improve detection performance, while failure analysis identifies challenges including context-dependent interpretation, cultural understanding, and cross-linguistic sarcasm detection, providing directions for future research in multilingual cyberbullying detection.

网络欺凌多语言可解释性Hinglish

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。