arXiv:2410.05287cs.CLcs.AI2024-10被引 8

用跨平台数据提升英德双语仇恨言论检测效果

Hate Speech Detection Using Cross-Platform Social Media Data In English and German Language

  • 融合YouTube、Twitter、Gab多平台评论数据训练模型
  • 英德语境下F1得分最高达0.74和0.68
  • 内容相似性与共用仇恨词是提升性能关键

仇恨言论已成为普遍现象,尤其在危机、选举和社会动荡时期加剧。尽管已有多种基于人工智能的检测方法,但通用模型仍未实现。文本分类中的主要挑战在于高质量训练数据的获取成本。本研究聚焦于检测英德双语环境下YouTube评论中的仇恨言论,并评估引入其他平台数据对分类模型性能的影响。我们考察了跨平台额外训练数据的价值,同时分析内容相似性、定义相似性和共用仇恨词汇等因素对模型表现的影响。结果表明,基于内容相似性、仇恨词汇和定义相似性构建的更多相似数据集能有效提升模型性能。最佳表现来自整合YouTube、Twitter和Gab数据集,英、德语YouTube评论的F1分数分别达到0.74和0.68。

原文摘要 · Abstract (English)

Hate speech has grown into a pervasive phenomenon, intensifying during times of crisis, elections, and social unrest. Multiple approaches have been developed to detect hate speech using artificial intelligence, but a generalized model is yet unaccomplished. The challenge for hate speech detection as text classification is the cost of obtaining high-quality training data. This study focuses on detecting bilingual hate speech in YouTube comments and measuring the impact of using additional data from other platforms in the performance of the classification model. We examine the value of additional training datasets from cross-platforms for improving the performance of classification models. We also included factors such as content similarity, definition similarity, and common hate words to measure the impact of datasets on performance. Our findings show that adding more similar datasets based on content similarity, hate words, and definitions improves the performance of classification models. The best performance was obtained by combining datasets from YouTube comments, Twitter, and Gab with an F1-score of 0.74 and 0.68 for English and German YouTube comments.

仇恨言论多语言跨平台

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。