arXiv:2508.01712cs.CVcs.AI2025-08被引 4

构建细粒度仇恨视频数据集,助力精准识别恶意内容

HateClipSeg: A Segment-Level Annotated Dataset for Fine-Grained Hate Video Detection

  • 采用三阶段标注流程,实现段落级细粒度标注
  • 涵盖超1万1千个片段,五类攻击性内容及受害者标签
  • 提出三项新任务,推动多模态时序分析模型发展

由于多模态内容复杂且现有数据集缺乏细粒度标注,视频中仇恨言论检测仍具挑战。我们推出 HateClipSeg,一个大规模多模态数据集,包含超过 11,714 个视频段落的段落级标注,类别分为正常与五类攻击性:仇恨、侮辱、性相关、暴力、自残,并附明确目标受害者标签。通过三阶段标注流程,达到高一致性的标注结果(Krippendorff's alpha = 0.817)。我们设计三项评测任务:(1)剪裁后仇恨视频分类,(2)仇恨内容时空定位,(3)在线仇恨视频分类。实验表明当前模型存在显著差距,凸显对更先进多模态与时序感知方法的需求。HateClipSeg 数据集已公开于 https://github.com/Social-AI-Studio/HateClipSeg.git。

原文摘要 · Abstract (English)

Detecting hate speech in videos remains challenging due to the complexity of multimodal content and the lack of fine-grained annotations in existing datasets. We present HateClipSeg, a large-scale multimodal dataset with both video-level and segment-level annotations, comprising over 11,714 segments labeled as Normal or across five Offensive categories: Hateful, Insulting, Sexual, Violence, Self-Harm, along with explicit target victim labels. Our three-stage annotation process yields high inter-annotator agreement (Krippendorff's alpha = 0.817). We propose three tasks to benchmark performance: (1) Trimmed Hateful Video Classification, (2) Temporal Hateful Video Localization, and (3) Online Hateful Video Classification. Results highlight substantial gaps in current models, emphasizing the need for more sophisticated multimodal and temporally aware approaches. The HateClipSeg dataset are publicly available at https://github.com/Social-AI-Studio/HateClipSeg.git.

仇恨检测多模态视频分析数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。