arXiv:2608.19006cs.CL2026-08中稿 · WOAH 2026

检测仇恨言论时别泄露作者隐私,新方法可平衡二者。

Introducing the Privacy-HSD Trade-off: Hate Speech Detection, but not at the Cost of Privacy

论文配图:Introducing the Privacy-HSD Trade-off: Hate Speech Detection, but not at the Cost of Privacy
图 1 · 摘自论文原文
  • 提出隐私-仇恨言论检测权衡机制,避免系统暴露作者身份。
  • 实验证明现有方法易泄露隐私,新方法AgnoSpeech有效缓解此问题。
  • 适合关注在线安全与数据隐私的AI研究者与政策制定者。

仇恨言论是影响大量在线用户(尤其是青少年和少数群体)的现实威胁。尽管构建可靠且鲁棒的自动仇恨言论检测(HSD)系统至关重要,但必须同时兼顾个人隐私权。我们探讨了HSD与隐私的交集,发现现有系统可能在提升性能的同时无意中编码作者身份,从而威胁隐私。基于此,我们提出隐私-HSD权衡概念,强调需谨慎平衡。我们评估了一系列文本隐私化方法,以及新提出的领域特定的AgnoSpeech技术,结果表明在保护隐私的同时实现有效的仇恨言论检测虽具挑战性但可行。研究呼吁更多关注隐私与HSD之间的权衡,二者均对保障网络参与具有实际意义。

原文摘要 · Abstract (English)

Hate speech is a real and timely threat that affects a large portion of online users, especially youth and minority groups. While building reliable and robust automatic hate speech detection (HSD) systems is paramount, we argue that this must also be balanced with the individual right to privacy. Exploring the intersection of HSD and privacy, we demonstrate that HSD systems might unintentionally achieve performance at the cost of encoding authorship, posing a threat to privacy. Building on these findings, we establish the notion of a privacy-HSD trade-off, which demands a careful balance. We benchmark a series of text privatization methods, as well as our newly proposed domain-specific AgnoSpeech technique, showing that balancing privacy and HSD is difficult but feasible. The findings make a strong case for more research on the trade-offs between privacy and HSD, both of which have tangible implications for the safeguarding of online participation.

仇恨言论检测隐私保护文本隐私化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。