对话上下文能显著提升恶意言论检测准确率,内容特征比账号特征更重要。
Context Matters: Incorporating Target Awareness in Conversational Abusive Language Detection
- 利用父帖内容作为上下文增强回复的恶意判断
- 结合多种内容特征使检测效果提升明显
- 适合需要真实对话场景的社交平台安全系统
恶意语言检测在社交媒体治理中日益重要。现有研究多单独分析单条帖子,忽略了对话中的上下文信息。本文聚焦于用户对前序帖子(父帖)的回复,探讨是否借助父帖上下文能更好判断回复是否为恶意内容。我们提取了基于内容和账户的多种上下文特征,并与仅使用回复自身特征的方法对比。在包含父-回复对的对话数据集上,测试四种分类模型。实验表明,引入上下文特征可显著提升性能;其中内容特征(如文本语义)比账户特征(如用户身份)贡献更大。综合运用多种内容特征优于单一或少量特征选择。研究为真实对话场景下的上下文感知恶意语言检测提供了实用洞见。
原文摘要 · Abstract (English)
Abusive language detection has become an increasingly important task as a means to tackle this type of harmful content in social media. There has been a substantial body of research developing models for determining if a social media post is abusive or not; however, this research has primarily focused on exploiting social media posts individually, overlooking additional context that can be derived from surrounding posts. In this study, we look at conversational exchanges, where a user replies to an earlier post by another user (the parent tweet). We ask: does leveraging context from the parent tweet help determine if a reply post is abusive or not, and what are the features that contribute the most? We study a range of content-based and account-based features derived from the context, and compare this to the more widely studied approach of only looking at the features from the reply tweet. For a more generalizable study, we test four different classification models on a dataset made of conversational exchanges (parent-reply tweet pairs) with replies labeled as abusive or not. Our experiments show that incorporating contextual features leads to substantial improvements compared to the use of features derived from the reply tweet only, confirming the importance of leveraging context. We observe that, among the features under study, it is especially the content-based features (what is being posted) that contribute to the classification performance rather than account-based features (who is posting it). While using content-based features, it is best to combine a range of different features to ensure improved performance over being more selective and using fewer features. Our study provides insights into the development of contextualized abusive language detection models in realistic settings involving conversations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。