构建尼日利亚音乐人社交媒体辱骂语言数据集,提升多语种检测能力
VocalTweets: Investigating Social Media Offensive Language Among Nigerian Musicians
- 采集12位尼日利亚音乐人推文,构建跨语言多语种辱骂文本数据集
- 基于Twitter-RoBERTa模型实现74.5的F1分数,有效识别辱骂内容
- 适用于社交媒体内容安全、文化语境下的语言分析研究者
音乐人常通过社交媒体表达观点,但其线上言论与音乐作品中的信息常不一致。部分人借平台攻击同行,也有人支持政见或参与社会运动,如#EndSars抗议。尽管已有大量关于社交媒体辱骂语言检测的研究,但针对音乐人使用辱骂语言的研究仍较少。本研究提出VocalTweets,一个包含12位知名尼日利亚音乐人推文的代码混杂、多语种数据集,采用二分类标注(正常/辱骂)。我们使用HuggingFace的base-Twitter-RoBERTa模型进行训练,获得74.5的F1分数。此外,通过在OLID数据集上进行跨语料库实验,评估了该数据集的泛化能力。
原文摘要 · Abstract (English)
Musicians frequently use social media to express their opinions, but they often convey different messages in their music compared to their posts online. Some utilize these platforms to abuse their colleagues, while others use it to show support for political candidates or engage in activism, as seen during the #EndSars protest. There are extensive research done on offensive language detection on social media, the usage of offensive language by musicians has received limited attention. In this study, we introduce VocalTweets, a code-switched and multilingual dataset comprising tweets from 12 prominent Nigerian musicians, labeled with a binary classification method as Normal or Offensive. We trained a model using HuggingFace's base-Twitter-RoBERTa, achieving an F1 score of 74.5. Additionally, we conducted cross-corpus experiments with the OLID dataset to evaluate the generalizability of our dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。