arXiv:2409.17191cs.CLcs.LG2024-09被引 5

新框架提升仇恨言论检测的准确率、鲁棒性和公平性

An Effective, Robust and Fairness-aware Hate Speech Detection Framework

  • 用双向四元数拟LSTM提升模型效率与效果
  • 融合五个平台数据集,性能超越8个先进方法
  • 兼顾公平性与不确定性估计,适合实际部署

随着在线社交网络的普及,仇恨言论传播更快、危害更大。现有检测方法在数据不足、模型不确定性估计、抗恶意攻击能力以及无意偏差(即公平性)方面存在局限。亟需在社交网络中实现准确、鲁棒且公平的仇恨言论分类。为此,我们设计了一种数据增强、公平性考量和不确定性估计的新框架。框架中提出双向四元数拟LSTM层,在效果与效率间取得平衡。为构建泛化能力强的模型,我们整合了来自三个平台的五个数据集。实验表明,该模型在无攻击和多种攻击场景下均优于八个最先进方法,验证了其有效性与鲁棒性。代码与合并数据集已公开,以促进后续研究。

原文摘要 · Abstract (English)

With the widespread online social networks, hate speeches are spreading faster and causing more damage than ever before. Existing hate speech detection methods have limitations in several aspects, such as handling data insufficiency, estimating model uncertainty, improving robustness against malicious attacks, and handling unintended bias (i.e., fairness). There is an urgent need for accurate, robust, and fair hate speech classification in online social networks. To bridge the gap, we design a data-augmented, fairness addressed, and uncertainty estimated novel framework. As parts of the framework, we propose Bidirectional Quaternion-Quasi-LSTM layers to balance effectiveness and efficiency. To build a generalized model, we combine five datasets collected from three platforms. Experiment results show that our model outperforms eight state-of-the-art methods under both no attack scenario and various attack scenarios, indicating the effectiveness and robustness of our model. We share our code along with combined dataset for better future research

仇恨言论检测公平性鲁棒性深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。