构建跨平台暴力威胁数据集,验证了模型在不同平台间通用性。
Cross-Platform Violence Detection on Social Media: A Dataset and Analysis
- 收集3万条跨平台暴力言论,标注政治与性暴力等子类型。
- 跨数据集训练测试准确率高,表明暴力特征具有普适性。
- 适合内容安全研究者、平台风控团队参考使用。
暴力威胁在社交媒体上仍是重大问题。高质量数据有助于理解与检测恶意内容。本文构建了一个包含3万条手标注暴力威胁的跨平台数据集,涵盖政治暴力与性暴力等子类型。通过与YouTube已有暴力评论数据集对比分析,发现即使来源平台和标注标准不同,模型在单数据集训练/跨数据集测试及合并数据集条件下均取得高分类准确率。结果对内容分类策略及跨平台暴力内容理解具有重要意义。
原文摘要 · Abstract (English)
Violent threats remain a significant problem across social media platforms. Useful, high-quality data facilitates research into the understanding and detection of malicious content, including violence. In this paper, we introduce a cross-platform dataset of 30,000 posts hand-coded for violent threats and sub-types of violence, including political and sexual violence. To evaluate the signal present in this dataset, we perform a machine learning analysis with an existing dataset of violent comments from YouTube. We find that, despite originating from different platforms and using different coding criteria, we achieve high classification accuracy both by training on one dataset and testing on the other, and in a merged dataset condition. These results have implications for content-classification strategies and for understanding violent content across social media.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。