arXiv:2508.06570cs.CVcs.LG2025-08ACL被引 19

构建首个大规模隐性仇恨言论视频数据集并提出两阶段对比学习框架。

ImpliHateVid: A Benchmark Dataset and Two-stage Contrastive Learning Framework for Implicit Hate Speech Detection in Videos

  • 分两阶段用对比学习融合音视频文本特征,提升多模态理解能力。
  • 在2009个视频上验证,隐性仇恨检测准确率显著优于基线模型。
  • 适合研究多模态内容安全、隐性暴力识别的学者与工程师。

现有研究主要集中在文本和图像层面的仇恨言论检测,而视频层面的方法仍不充分。本文提出一个新数据集ImpliHateVid,专为视频中隐性仇恨言论检测设计,包含2,009个视频:509个隐性仇恨视频、500个显性仇恨视频和1,000个非仇恨视频,是首个大规模专注于隐性仇恨检测的视频数据集。我们还提出一种两阶段对比学习框架:第一阶段使用对比损失训练音频、文本、图像的模态专用编码器,并拼接三者特征;第二阶段通过对比学习训练跨模态编码器,优化多模态表示。同时引入情感、情绪及字幕特征以增强隐性仇恨识别。在两个数据集上评估:ImpliHateVid用于隐性仇恨检测,HateMM用于通用仇恨言论检测,结果表明该多模态对比学习方法有效,且数据集具有重要价值。

原文摘要 · Abstract (English)

The existing research has primarily focused on text and image-based hate speech detection, video-based approaches remain underexplored. In this work, we introduce a novel dataset, ImpliHateVid, specifically curated for implicit hate speech detection in videos. ImpliHateVid consists of 2,009 videos comprising 509 implicit hate videos, 500 explicit hate videos, and 1,000 non-hate videos, making it one of the first large-scale video datasets dedicated to implicit hate detection. We also propose a novel two-stage contrastive learning framework for hate speech detection in videos. In the first stage, we train modality-specific encoders for audio, text, and image using contrastive loss by concatenating features from the three encoders. In the second stage, we train cross-encoders using contrastive learning to refine multimodal representations. Additionally, we incorporate sentiment, emotion, and caption-based features to enhance implicit hate detection. We evaluate our method on two datasets, ImpliHateVid for implicit hate speech detection and another dataset for general hate speech detection in videos, HateMM dataset, demonstrating the effectiveness of the proposed multimodal contrastive learning for hateful content detection in videos and the significance of our dataset.

视频检测多模态仇恨言论对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。