综述社交平台识别性侵者文本的方法与挑战
A Survey on Pedophile Attribution Techniques for Online Platforms
- 系统梳理用于识别性侵者文本的算法与特征
- 发现现有方法无法实现嫌疑人的准确归属
- 适合安全、内容审核方向的研究者参考
社交媒体中对匿名性的依赖使其广受欢迎,公共Wi-Fi的普及也促进了各类在线内容的传播。尽管匿名和易访问性为用户提供了便利,但难以保护弱势群体免受性侵害者威胁。开发能将文本与嫌疑人关联的自动化系统可提升应对能力。本文综述了社交平台上用于性侵者归属的技术方法,分析了嫌疑人群体规模与文本长度对归属任务的影响。我们还回顾了常用数据集、特征、分类技术及评估指标。研究发现,虽有少量研究提出缓解在线性侵风险的工具,但均无法实现嫌疑人归属。最后,本文列出了若干开放性研究问题。
原文摘要 · Abstract (English)
Reliance on anonymity in social media has increased its popularity on these platforms among all ages. The availability of public Wi-Fi networks has facilitated a vast variety of online content, including social media applications. Although anonymity and ease of access can be a convenient means of communication for their users, it is difficult to manage and protect its vulnerable users against sexual predators. Using an automated identification system that can attribute predators to their text would make the solution more attainable. In this survey, we provide a review of the methods of pedophile attribution used in social media platforms. We examine the effect of the size of the suspect set and the length of the text on the task of attribution. Moreover, we review the most-used datasets, features, classification techniques and performance measures for attributing sexual predators. We found that few studies have proposed tools to mitigate the risk of online sexual predators, but none of them can provide suspect attribution. Finally, we list several open research problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。