用大模型生成数据和标签,让儿童网络欺凌检测更高效安全
Synthetic vs. Gold: The Role of LLM Generated Labels and Data in Cyberbullying Detection
- 用大模型生成儿童语言风格的合成数据和标签
- 合成数据训练的模型准确率达75.8%,接近真实数据的81.5%
- 适合需要快速构建伦理合规检测系统的研究者
网络欺凌对儿童构成严重威胁,亟需高效检测系统保障在线安全。现有大规模辱骂数据集虽存在,但缺乏反映儿童语言风格的标注数据。从儿童群体获取真实数据面临伦理、法律与技术障碍,且人工标注耗时耗力,还可能使标注者暴露于有害内容。本文利用大语言模型(LLMs)生成合成数据与标签,实验表明:基于合成数据训练的BERT分类器准确率达75.8%,接近使用完全真实数据的81.5%;同时,LLMs能有效为真实未标注数据打标,使分类器性能达79.1%,与真实数据训练结果相近。结果证明,大模型可作为可扩展、伦理友好且低成本的解决方案,用于构建网络欺凌检测数据。
原文摘要 · Abstract (English)
Cyberbullying (CB) presents a pressing threat, especially to children, underscoring the urgent need for robust detection systems to ensure online safety. While large-scale datasets on online abuse exist, there remains a significant gap in labeled data that specifically reflects the language and communication styles used by children. The acquisition of such data from vulnerable populations, such as children, is challenging due to ethical, legal and technical barriers. Moreover, the creation of these datasets relies heavily on human annotation, which not only strains resources but also raises significant concerns due to annotators exposure to harmful content. In this paper, we address these challenges by leveraging Large Language Models (LLMs) to generate synthetic data and labels. Our experiments demonstrate that synthetic data enables BERT-based CB classifiers to achieve performance close to that of those trained on fully authentic datasets (75.8% vs. 81.5% accuracy). Additionally, LLMs can effectively label authentic yet unlabeled data, allowing BERT classifiers to attain a comparable performance level (79.1% vs. 81.5% accuracy). These results highlight the potential of LLMs as a scalable, ethical, and cost-effective solution for generating data for CB detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。