用多模态专家模型提升中文仇恨言论检测抗伪装能力
MMBERT: Scaled Mixture-of-Experts Multimodal BERT for Robust Chinese Hate Speech Detection under Cloaking Perturbations
- 融合文本、语音、视觉的专家混合架构,动态分配模态处理路径
- 在多个中文仇恨言论数据集上超越传统BERT与大模型基线
- 特别适合对抗伪装攻击的中文内容安全场景
中文社交网络中的仇恨言论检测面临独特挑战,尤其因广泛使用逃避文本检测系统的伪装技术。尽管大语言模型(LLMs)近期提升了检测能力,但多数研究集中于英文数据集,且对中文语境下的多模态策略关注有限。本文提出MMBERT,一种基于BERT的多模态框架,通过混合专家(MoE)架构整合文本、语音和视觉模态。为解决直接将MoE引入BERT模型带来的不稳定性,我们设计了渐进式三阶段训练范式。MMBERT包含模态专属专家、共享自注意力机制及基于路由器的专家分配策略,显著增强对抗对抗性扰动的鲁棒性。在多个中文仇恨言论数据集上的实证结果表明,MMBERT显著优于微调后的BERT编码器模型、微调后的LLMs以及采用上下文学习的大模型方法。
原文摘要 · Abstract (English)
Hate speech detection on Chinese social networks presents distinct challenges, particularly due to the widespread use of cloaking techniques designed to evade conventional text-based detection systems. Although large language models (LLMs) have recently improved hate speech detection capabilities, the majority of existing work has concentrated on English datasets, with limited attention given to multimodal strategies in the Chinese context. In this study, we propose MMBERT, a novel BERT-based multimodal framework that integrates textual, speech, and visual modalities through a Mixture-of-Experts (MoE) architecture. To address the instability associated with directly integrating MoE into BERT-based models, we develop a progressive three-stage training paradigm. MMBERT incorporates modality-specific experts, a shared self-attention mechanism, and a router-based expert allocation strategy to enhance robustness against adversarial perturbations. Empirical results in several Chinese hate speech datasets show that MMBERT significantly surpasses fine-tuned BERT-based encoder models, fine-tuned LLMs, and LLMs utilizing in-context learning approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。