arXiv:2503.06534cs.CL2025-03NAACL被引 4

SafeSpeech可检测对话中隐性性别歧视与暴力语言,支持多粒度分析。

SafeSpeech: A Comprehensive and Interactive Tool for Analysing Sexist and Abusive Language in Conversations

  • 融合微调分类器与大模型,实现消息与对话层面的联合检测
  • 在EDOS、OffensEval等数据集上达到顶尖性能,可识别细粒度性别偏见
  • 提供可解释性分析,适合研究网络暴力与内容安全的学者使用

检测包括性别歧视、骚扰和攻击性行为在内的毒性语言仍是重大挑战,尤其在隐含且依赖上下文的情况下。现有方法多聚焦于孤立消息级别的分类,忽略了跨对话语境中涌现的毒性。为推动该方向的研究,我们提出SafeSpeech——一个综合性平台,连接消息级与对话级洞察。该平台集成微调分类器与大语言模型(LLMs),支持多粒度检测、毒性感知对话摘要及角色画像生成。SafeSpeech还引入可解释性机制,如困惑度提升分析,以突出驱动预测的语言要素。在EDOS、OffensEval和HatEval等基准数据集上的评估表明,其在多个任务上复现了当前最优性能,包括细粒度性别歧视检测。

原文摘要 · Abstract (English)

Detecting toxic language including sexism, harassment and abusive behaviour, remains a critical challenge, particularly in its subtle and context-dependent forms. Existing approaches largely focus on isolated message-level classification, overlooking toxicity that emerges across conversational contexts. To promote and enable future research in this direction, we introduce SafeSpeech, a comprehensive platform for toxic content detection and analysis that bridges message-level and conversation-level insights. The platform integrates fine-tuned classifiers and large language models (LLMs) to enable multi-granularity detection, toxic-aware conversation summarization, and persona profiling. SafeSpeech also incorporates explainability mechanisms, such as perplexity gain analysis, to highlight the linguistic elements driving predictions. Evaluations on benchmark datasets, including EDOS, OffensEval, and HatEval, demonstrate the reproduction of state-of-the-art performance across multiple tasks, including fine-grained sexism detection.

毒性检测对话分析可解释性性别歧视

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。