研究表情符号如何在社交媒体中助长伤害性内容并提出智能过滤方案
The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation
- 用大模型构建多阶段审核流程,精准替换有害表情符号
- 人类评估显示可降低攻击性感知,同时保留原意
- 揭示不同类型冒犯内容对表情符号的依赖差异
社交媒体已成为现代沟通的核心,但也充斥着挑战平台安全与包容性的冒犯性内容。尽管以往研究多关注文本层面的冒犯指标,但表情符号——这种在线对话中普遍存在的视觉元素——的作用仍被忽视。尽管单个表情符号很少具有冒犯性,但通过象征关联、反讽或语境误用,它们可能产生有害含义。本文系统分析表情符号在推特冒犯性消息中的贡献,考察其在不同冒犯类别中的分布及用户如何利用表情符号的模糊性。为此,我们提出一种基于大模型的多步骤内容审核流程,能选择性替换有害表情符号,同时保持推文语义意图。人工评估证实该方法有效降低内容冒犯感,且不损失原意。分析还揭示了不同冒犯类型间效果的异质性,为在线交流与表情符号治理提供细致洞见。
原文摘要 · Abstract (English)
Social media platforms have become central to modern communication, yet they also harbor offensive content that challenges platform safety and inclusivity. While prior research has primarily focused on textual indicators of offense, the role of emojis, ubiquitous visual elements in online discourse, remains underexplored. Emojis, despite being rarely offensive in isolation, can acquire harmful meanings through symbolic associations, sarcasm, and contextual misuse. In this work, we systematically examine emoji contributions to offensive Twitter messages, analyzing their distribution across offense categories and how users exploit emoji ambiguity. To address this, we propose an LLM-powered, multi-step moderation pipeline that selectively replaces harmful emojis while preserving the tweet's semantic intent. Human evaluations confirm our approach effectively reduces perceived offensiveness without sacrificing meaning. Our analysis also reveals heterogeneous effects across offense types, offering nuanced insights for online communication and emoji moderation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。