用Hinglish版Transformer自动分类网络犯罪举报,提升执法效率。
Automated Classification of Cybercrime Complaints using Transformer-based Language Models for Hinglish Texts
- 采用适配印地语-英语混杂文本的HingBERT/HingRoBERTa模型处理复杂语言
- 在真实数据集上实现74.41%准确率和71.49%的F1分数
- 兼顾隐私保护与可部署性,适合公安与网安平台使用
网络犯罪上升及多语言、代码混杂举报文本的复杂性给执法与网络安全机构带来挑战。这些机构需要自动化、可扩展的方法来识别罪行类型,以高效处理大量举报。人工分类效率低,传统机器学习难以捕捉文本的语义和上下文特征。此外,缺乏公开数据集和隐私问题阻碍了稳健解决方案的研究。为此,我们提出一个自动化网络犯罪举报分类框架。该框架利用专为Hinglish设计的Transformer模型(如HingBERT和HingRoBERTa)有效处理代码混杂输入。我们使用印度网络犯罪协调中心(I4C)在2024年CyberGuard AI黑客松提供的真实数据集。通过基于开源GenAI模型的数据增强方法缓解类别不平衡问题,并采用隐私友好的预处理流程,在保障伦理标准的同时维护数据完整性。我们的方案取得显著性能提升:HingRoBERTa达到74.41%准确率和71.49% F1分数。我们还开发了一个集成Django REST后端与现代前端的即用型工具,具备可扩展性,可直接部署于国家网络犯罪举报门户等平台。本工作填补了网络犯罪举报管理中的关键空白,提供了一种可扩展、注重隐私且适应性强的现代网络安全解决方案。
原文摘要 · Abstract (English)
The rise in cybercrime and the complexity of multilingual and code-mixed complaints present significant challenges for law enforcement and cybersecurity agencies. These organizations need automated, scalable methods to identify crime types, enabling efficient processing and prioritization of large complaint volumes. Manual triaging is inefficient, and traditional machine learning methods fail to capture the semantic and contextual nuances of textual cybercrime complaints. Moreover, the lack of publicly available datasets and privacy concerns hinder the research to present robust solutions. To address these challenges, we propose a framework for automated cybercrime complaint classification. The framework leverages Hinglish-adapted transformers, such as HingBERT and HingRoBERTa, to handle code-mixed inputs effectively. We employ the real-world dataset provided by Indian Cybercrime Coordination Centre (I4C) during CyberGuard AI Hackathon 2024. We employ GenAI open source model-based data augmentation method to address class imbalance. We also employ privacy-aware preprocessing to ensure compliance with ethical standards while maintaining data integrity. Our solution achieves significant performance improvements, with HingRoBERTa attaining an accuracy of 74.41% and an F1-score of 71.49%. We also develop ready-to-use tool by integrating Django REST backend with a modern frontend. The developed tool is scalable and ready for real-world deployment in platforms like the National Cyber Crime Reporting Portal. This work bridges critical gaps in cybercrime complaint management, offering a scalable, privacy-conscious, and adaptable solution for modern cybersecurity challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。