破解多模态仇恨言论中隐含意图的演变,提升对隐蔽攻击的识别能力。
More Than Sum of Its Parts: Deciphering Intent Shifts in Multimodal Hate Speech Detection
- 通过模拟法庭辩论机制,让模型在多模态交互中深入分析隐含语义。
- 在新构建的H-VLI数据集上,对隐性仇恨言论的准确率显著优于现有方法。
- 适合关注多模态内容安全、情感与意图理解的研究者使用。
社交媒体中的仇恨言论治理至关重要,但依赖于自动化检测系统的有效性。随着内容形式演进,仇恨言论正从纯文本转向复杂的多模态表达,使得隐含攻击更难察觉。现有系统在这些微妙情形下表现不佳,因难以处理模态间相互作用产生的超越单个模态叠加意义的新型语义。为此,本文不再局限于二分类任务,而是聚焦于模态交互导致的语义意图转变——即原本无害的线索组合成隐性仇恨,或通过语义反转消解毒性。基于此细粒度定义,我们构建了「仇恨通过视觉-语言互动」(Hate via Vision-Language Interplay, H-VLI)基准数据集,其真实意图取决于模态间的复杂互动而非明显的视觉或文本辱骂。为有效解析此类复杂线索,我们提出「不对称推理通过法庭代理辩论」(ARCADE)框架,模拟司法审判过程,让代表控诉与辩护的代理主动交锋,迫使模型在作出判断前深度审视深层语义。大量实验表明,ARCADE在H-VLI上显著优于当前最优基线,尤其在挑战性的隐性案例上表现突出,同时在传统基准上保持竞争力。代码与数据已公开于:https://github.com/Sayur1n/H-VLI
原文摘要 · Abstract (English)
Combating hate speech on social media is critical for securing cyberspace, yet relies heavily on the efficacy of automated detection systems. As content formats evolve, hate speech is transitioning from solely plain text to complex multimodal expressions, making implicit attacks harder to spot. Current systems, however, often falter on these subtle cases, as they struggle with multimodal content where the emergent meaning transcends the aggregation of individual modalities. To bridge this gap, we move beyond binary classification to characterize semantic intent shifts where modalities interact to construct implicit hate from benign cues or neutralize toxicity through semantic inversion. Guided by this fine-grained formulation, we curate the Hate via Vision-Language Interplay (H-VLI) benchmark where the true intent hinges on the intricate interplay of modalities rather than overt visual or textual slurs. To effectively decipher these complex cues, we further propose the Asymmetric Reasoning via Courtroom Agent DEbate (ARCADE) framework. By simulating a judicial process where agents actively argue for accusation and defense, ARCADE forces the model to scrutinize deep semantic cues before reaching a verdict. Extensive experiments demonstrate that ARCADE significantly outperforms state-of-the-art baselines on H-VLI, particularly for challenging implicit cases, while maintaining competitive performance on established benchmarks. Our code and data are available at: https://github.com/Sayur1n/H-VLI
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。