通过细粒度片段分析,提升仇恨视频检测精度
CLARA: Clip-Level Multimodal Alignment with VLM-Derived Rationales for Hateful Video Detection

- 将视频拆分为片段,用多模态专家协作捕捉局部仇恨信号
- 在三个数据集上优于当前最优方法,显著提升检测准确率
- 适合需要高精度视频内容安全审核的研究与应用
随着以视频为主的社会媒体平台迅速发展,仇恨言论对个人福祉与社会凝聚力构成严重威胁,仇恨视频检测日益重要。相较于文本或静态多模态内容,仇恨视频检测研究较少且更具挑战性,因仇恨含义常源于语音、音频与视觉等多模态线索的复杂交互,且信号短暂、隐含且具有时间依赖性,传统视频级表示难以捕捉。本文提出CLARA,一种基于片段级别的多模态框架用于仇恨视频检测。不同于将视频视为单一整体,CLARA将其建模为细粒度片段序列,更精准定位时序上的仇恨信号。引入混合专家片段编码器实现自适应多模态对齐,设计局部-全局段对比目标联合建模短时线索与长时依赖,并通过门控Transformer融合视觉语言模型生成的推理依据,提供高层语义引导。在三个仇恨视频数据集上的大量实验表明,CLARA持续优于现有最先进方法。消融实验与参数分析验证了各组件的有效性。
原文摘要 · Abstract (English)
Hateful video detection has become increasingly important with the rapid growth of video-centric social media platforms, given the serious risks that hate speech poses to both individual well-being and social cohesion. Compared with text or static multimodal content, hateful video detection remains underexplored and significantly more challenging, as hateful meaning often arises from complex interactions among multimodal cues, including speech, audio, and visual content. Moreover, such signals are often brief, implicit, and temporally dependent, making them difficult to capture using conventional video-level representations. In this work, we propose CLARA, a clip-level multimodal framework for hateful video detection. Instead of treating a video as a single instance, CLARA models it as a sequence of fine-grained clips, enabling more precise capture of temporally localized hateful signals. We introduce a Mixture-of-Experts clip encoder for adaptive multimodal alignment, a local-global segment contrastive objective to jointly model short-term cues and long-range temporal dependencies, and VLM-derived rationales integrated via a gated Transformer to provide high-level semantic guidance. Extensive experiments on three hateful video datasets demonstrate that CLARA consistently outperforms state-of-the-art methods. Further ablation studies and parameter analyses validate the effectiveness of each component.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。