arXiv:2508.20546cs.MMcs.AI2025-08中稿 · ACM Multimedia 202…被引 11

融合视频、音频、字幕的多模态仇恨言论检测模型

MM-HSD: Multi-Modal Hate Speech Detection in Videos

  • 用跨模态注意力早期提取特征,融合视频、音频和字幕信息
  • 在HateMM数据集上达到0.874的M-F1,优于现有方法
  • 首次系统比较查询/键配置,适合多模态内容安全研究者

尽管仇恨言论检测(HSD)在文本领域已广泛研究,现有视频多模态方法仍受限,尤其当各模态单独信息量不足时,简单融合无法充分捕捉模态间依赖。此外,以往研究常忽略屏幕字幕和音频等关键模态,而这些可能包含细微仇恨内容,对判断至关重要。本文提出MM-HSD,一种整合视频帧、音频、语音转录文本及屏幕字幕的多模态仇恨言论检测模型,并引入跨模态注意力(CMA)作为早期特征提取器。我们首次系统比较不同查询/键配置,并评估模态在CMA模块中的交互作用。实验表明,以屏幕字幕为查询、其余模态为键时性能最优。在HateMM数据集上,使用转录文本、音频、视频、屏幕字幕与CMA特征拼接,原始模态嵌入下取得0.874的M-F1,优于当前最优方法。代码已开源。

原文摘要 · Abstract (English)

While hate speech detection (HSD) has been extensively studied in text, existing multi-modal approaches remain limited, particularly in videos. As modalities are not always individually informative, simple fusion methods fail to fully capture inter-modal dependencies. Moreover, previous work often omits relevant modalities such as on-screen text and audio, which may contain subtle hateful content and thus provide essential cues, both individually and in combination with others. In this paper, we present MM-HSD, a multi-modal model for HSD in videos that integrates video frames, audio, and text derived from speech transcripts and from frames (i.e.~on-screen text) together with features extracted by Cross-Modal Attention (CMA). We are the first to use CMA as an early feature extractor for HSD in videos, to systematically compare query/key configurations, and to evaluate the interactions between different modalities in the CMA block. Our approach leads to improved performance when on-screen text is used as a query and the rest of the modalities serve as a key. Experiments on the HateMM dataset show that MM-HSD outperforms state-of-the-art methods on M-F1 score (0.874), using concatenation of transcript, audio, video, on-screen text, and CMA for feature extraction on raw embeddings of the modalities. The code is available at https://github.com/idiap/mm-hsd

多模态仇恨言论视频分析跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。