TANDEM让仇恨言论检测能定位时间点并识别目标,比现有方法提升30%准确率。
TANDEM: Temporal-Aware Neural Detection for Multimodal Hate Speech
- 采用视听与文本模型互训机制,实现跨模态自约束推理。
- 在HateMM数据集上目标识别F1达0.73,较当前最优提升30%。
- 适合需要可解释性、精准定位的平台内容审核场景。
社交媒体中长视频多模态内容日益增多,有害叙事通过音频、视觉和文本线索复杂交织。现有自动检测系统虽准确率高,但缺乏细粒度可解释证据,如精确时间戳与目标身份,难以支持人工介入审核。本文提出TANDEM框架,将多模态仇恨言论检测从二分类任务转化为结构化推理问题。该方法采用新颖的协同强化学习策略,使视觉-语言与音频-语言模型通过自约束跨模态上下文相互优化,在无需密集帧级标注的情况下稳定处理长时序序列。在三个基准数据集上的实验表明,TANDEM显著优于零样本及上下文增强基线,在HateMM数据集上目标识别F1达到0.73(较当前最优提升30%),同时保持精确的时间定位能力。进一步发现,尽管二分类检测表现稳健,但在多类别设置下区分攻击性与仇恨内容仍具挑战,源于标签模糊性和数据集不平衡。研究结果表明,即使在复杂多模态场景中,也可实现结构化、可解释的对齐,为下一代透明且可操作的在线安全审核工具提供范本。
原文摘要 · Abstract (English)
Social media platforms are increasingly dominated by long-form multimodal content, where harmful narratives are constructed through a complex interplay of audio, visual, and textual cues. While automated systems can flag hate speech with high accuracy, they often function as "black boxes" that fail to provide the granular, interpretable evidence, such as precise timestamps and target identities, required for effective human-in-the-loop moderation. In this work, we introduce TANDEM, a unified framework that transforms audio-visual hate detection from a binary classification task into a structured reasoning problem. Our approach employs a novel tandem reinforcement learning strategy where vision-language and audio-language models optimize each other through self-constrained cross-modal context, stabilizing reasoning over extended temporal sequences without requiring dense frame-level supervision. Experiments across three benchmark datasets demonstrate that TANDEM significantly outperforms zero-shot and context-augmented baselines, achieving 0.73 F1 in target identification on HateMM (a 30% improvement over state-of-the-art) while maintaining precise temporal grounding. We further observe that while binary detection is robust, differentiating between offensive and hateful content remains challenging in multi-class settings due to inherent label ambiguity and dataset imbalance. More broadly, our findings suggest that structured, interpretable alignment is achievable even in complex multimodal settings, offering a blueprint for the next generation of transparent and actionable online safety moderation tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。