arXiv:2410.20170cs.SIcs.LG2024-10被引 2

用AI区分网络讽刺与霸凌,准确率达95.15%。

Cyberbullying or just Sarcasm? Unmasking Coordinated Networks on Reddit

  • 结合NLP与机器学习构建区分框架
  • 在Reddit数据上达到95.15%识别准确率
  • 揭示青少年和少数群体易受攻击及协同霸凌模式

随着社交媒体使用量激增,用户常在帖子下发表讽刺性评论。虽然讽刺有时无害,但在负面或有害语境中可能与网络霸凌混淆。互联网的匿名性和广泛传播性加剧了这一问题,使平台如Reddit上的网络霸凌成为重大隐患。本研究聚焦于区分网络霸凌与讽刺,尤其针对在线语言细微差别导致难以判断恶意意图的难题。提出一种融合自然语言处理(NLP)与机器学习的框架,克服传统情感分析在检测复杂行为上的局限。通过分析从Reddit抓取的自定义数据集,实现了95.15%的准确率,能有效区分有害内容与讽刺。研究还发现青少年及少数群体尤为易受网络霸凌影响,并揭示了霸凌者背后的协同行为图谱,识别出其共同行为模式。该研究有助于提升在线社区的安全检测能力。

原文摘要 · Abstract (English)

With the rapid growth of social media usage, a common trend has emerged where users often make sarcastic comments on posts. While sarcasm can sometimes be harmless, it can blur the line with cyberbullying, especially when used in negative or harmful contexts. This growing issue has been exacerbated by the anonymity and vast reach of the internet, making cyberbullying a significant concern on platforms like Reddit. Our research focuses on distinguishing cyberbullying from sarcasm, particularly where online language nuances make it difficult to discern harmful intent. This study proposes a framework using natural language processing (NLP) and machine learning to differentiate between the two, addressing the limitations of traditional sentiment analysis in detecting nuanced behaviors. By analyzing a custom dataset scraped from Reddit, we achieved a 95.15% accuracy in distinguishing harmful content from sarcasm. Our findings also reveal that teenagers and minority groups are particularly vulnerable to cyberbullying. Additionally, our research uncovers coordinated graphs of groups involved in cyberbullying, identifying common patterns in their behavior. This research contributes to improving detection capabilities for safer online communities.

网络霸凌讽刺识别NLP社交网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。