arXiv:2507.04350cs.CL2025-07ACL被引 3

整合政策、平台与研究,构建统一仇恨言论自动治理框架

HatePRISM: Policies, Platforms, and Research Integration. Advancing NLP for Hate Speech Proactive Mitigation

  • 从国家法规、平台政策和研究数据三方面系统分析仇恨言论治理
  • 发现各国各平台对仇恨言论定义与处理方式差异显著
  • 提出融合主动净化与反言辞策略的统一治理研究方向

尽管各国及社交媒体平台已出台相关监管措施(如印度政府2021年规定、欧盟议会与理事会2022年指令),仇恨内容仍构成重大挑战。现有方法多依赖事后封禁或停用等被动手段,新兴策略则聚焦于主动净化与反言辞干预。本文提出的HatePRISM框架,从国家法规、平台政策与自然语言处理研究数据集三个维度,全面审视仇恨言论治理现状。研究发现,不同司法管辖区与平台在仇恨言论定义与内容审核实践上存在显著不一致,且与学术研究缺乏协同。基于此,我们提出构建融合多种策略的自动化仇恨言论治理统一框架的研究路径。

原文摘要 · Abstract (English)

Despite regulations imposed by nations and social media platforms, e.g. (Government of India, 2021; European Parliament and Council of the European Union, 2022), inter alia, hateful content persists as a significant challenge. Existing approaches primarily rely on reactive measures such as blocking or suspending offensive messages, with emerging strategies focusing on proactive measurements like detoxification and counterspeech. In our work, which we call HatePRISM, we conduct a comprehensive examination of hate speech regulations and strategies from three perspectives: country regulations, social platform policies, and NLP research datasets. Our findings reveal significant inconsistencies in hate speech definitions and moderation practices across jurisdictions and platforms, alongside a lack of alignment with research efforts. Based on these insights, we suggest ideas and research direction for further exploration of a unified framework for automated hate speech moderation incorporating diverse strategies.

仇恨言论政策研究NLP治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。