arXiv:2510.10971cs.CLcs.AI2025-10ACL

用强化学习动态调整模块权重,提升隐性仇恨言论检测准确率

RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection

  • 多模块分工处理不同语言特征,通过强化学习优化权重
  • 在多个数据集上优于传统静态方法,显著提升检测精度
  • 输出可解释结果,适合需要理解数据特性的研究者

仇恨言论在社会中持续存在且形式不断演变。互联网和在线匿名性的进步加速了其传播并增加了检测难度。然而,仇恨言论数据集因来源和平台不同而具有多样特征,反映不同的语言风格和社会背景。以往研究常采用固定方法,未能适配数据特性。我们提出RV-HATE框架,针对每个数据集的特异性设计多模块检测系统,各模块专注于特定语言或上下文特征。通过强化学习优化各模块贡献权重,并采用投票机制融合输出。该方法在提升检测准确性的同时,提供对数据集独特特征的可解释洞察。实验表明,该方法有效应对隐性仇恨言论,在多个数据集上优于传统静态方法。

原文摘要 · Abstract (English)

Hate speech remains prevalent in human society and continues to evolve in its forms and expressions. Modern advancements in internet and online anonymity accelerate its rapid spread and complicate its detection. However, hate speech datasets exhibit diverse characteristics primarily because they are constructed from different sources and platforms, each reflecting different linguistic styles and social contexts. Despite this diversity, prior studies on hate speech detection often rely on fixed methodologies without adapting to data-specific features. We introduce RV-HATE, a detection framework designed to account for the dataset-specific characteristics of each hate speech dataset. RV-HATE consists of multiple specialized modules, where each module focuses on distinct linguistic or contextual features of hate speech. The framework employs reinforcement learning to optimize weights that determine the contribution of each module for a given dataset. A voting mechanism then aggregates the module outputs to produce the final decision. RV-HATE offers two primary advantages: (1)~it improves detection accuracy by tailoring the detection process to dataset-specific attributes, and (2)~it also provides interpretable insights into the distinctive features of each dataset. Consequently, our approach effectively addresses implicit hate speech and achieves superior performance compared to conventional static methods. Our code is available at https://github.com/leeyejin1231/RV-HATE.

仇恨言论检测强化学习多模块融合可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。