arXiv:2512.10408cs.CV2025-12被引 3

首个弱监督多模态仇恨内容定位框架,精准识别视频中何时出现仇恨内容。

MultiHateLoc: Towards Temporal Localisation of Multimodal Hate Content in Online Videos

  • 设计模态感知时序编码器,捕捉视觉、音频、文本流的异构动态特征。
  • 在弱监督下实现帧级定位,准确率超越现有方法,在两个数据集上表现领先。
  • 适合内容安全团队用于自动化检测短视频中的仇恨言论,支持多模态分析。

TikTok 和 YouTube 等平台视频内容的快速增长加剧了多模态仇恨话语的传播,其有害信号在视觉、听觉和文本流中以微妙且异步的方式出现。现有研究主要关注视频级别分类,而对实际至关重要的时间定位任务——即识别仇恨片段发生的具体时刻——仍缺乏有效方法,尤其在仅提供视频级标签的弱监督条件下,传统静态融合或分类架构难以捕捉跨模态与时间动态。为此,我们提出 MultiHateLoc,首个专为弱监督多模态仇恨定位设计的框架。该框架包含:(1) 模态感知时序编码器,建模异构序列模式,并引入针对文本的预处理模块增强特征;(2) 动态跨模态融合机制,自适应强调每时刻最具信息量的模态,并采用跨模态对比对齐策略提升特征一致性;(3) 模态感知的 MIL 目标函数,基于视频级标签识别判别性片段。尽管仅依赖粗粒度标签,MultiHateLoc 仍能生成细粒度、可解释的帧级预测。在 HateMM 与 MultiHateClip 数据集上的实验表明,本方法在定位任务中达到当前最优性能。

原文摘要 · Abstract (English)

The rapid growth of video content on platforms such as TikTok and YouTube has intensified the spread of multimodal hate speech, where harmful cues emerge subtly and asynchronously across visual, acoustic, and textual streams. Existing research primarily focuses on video-level classification, leaving the practically crucial task of temporal localisation, identifying when hateful segments occur, largely unaddressed. This challenge is even more noticeable under weak supervision, where only video-level labels are available, and static fusion or classification-based architectures struggle to capture cross-modal and temporal dynamics. To address these challenges, we propose MultiHateLoc, the first framework designed for weakly-supervised multimodal hate localisation. MultiHateLoc incorporates (1) modality-aware temporal encoders to model heterogeneous sequential patterns, including a tailored text-based preprocessing module for feature enhancement; (2) dynamic cross-modal fusion to adaptively emphasise the most informative modality at each moment and a cross-modal contrastive alignment strategy to enhance multimodal feature consistency; (3) a modality-aware MIL objective to identify discriminative segments under video-level supervision. Despite relying solely on coarse labels, MultiHateLoc produces fine-grained, interpretable frame-level predictions. Experiments on HateMM and MultiHateClip show that our method achieves state-of-the-art performance in the localisation task.

多模态仇恨内容时间定位弱监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。