无需训练即可精确定位视频中的仇恨内容,靠大模型实现跨模态分析。
Towards Training-free Multimodal Hate Localisation with Large Language Models
- 用大模型和多阶段提示分解视频五种模态,不依赖标注数据
- 在HateMM和MultiHateClip上超越所有无训练基线,显著提升定位精度
- 适合需要快速部署、可解释性强的仇恨内容检测场景
在线视频中仇恨内容的泛滥对个人福祉和社会和谐构成严重威胁。现有视频仇恨检测方法要么高度依赖大规模人工标注,要么缺乏精细的时间定位能力。本文提出LELA,首个基于大语言模型(LLM)的免训练仇恨视频定位框架。与依赖监督流程的现有模型不同,LELA利用大模型和模态特定描述生成,在无需训练的情况下实现仇恨内容的检测与时间定位。方法将视频分解为图像、语音、OCR、音乐和视频上下文五种模态,通过多阶段提示机制为每帧计算细粒度仇恨得分,并引入组合匹配机制增强跨模态推理。在两个具有挑战性的基准数据集HateMM和MultiHateClip上的实验表明,LELA大幅优于所有现有无训练基线。我们还进行了广泛的消融实验和定性可视化,确立了LELA作为可扩展、可解释的仇恨视频定位强基线的地位。
原文摘要 · Abstract (English)
The proliferation of hateful content in online videos poses severe threats to individual well-being and societal harmony. However, existing solutions for video hate detection either rely heavily on large-scale human annotations or lack fine-grained temporal precision. In this work, we propose LELA, the first training-free Large Language Model (LLM) based framework for hate video localization. Distinct from state-of-the-art models that depend on supervised pipelines, LELA leverages LLMs and modality-specific captioning to detect and temporally localize hateful content in a training-free manner. Our method decomposes a video into five modalities, including image, speech, OCR, music, and video context, and uses a multi-stage prompting scheme to compute fine-grained hateful scores for each frame. We further introduce a composition matching mechanism to enhance cross-modal reasoning. Experiments on two challenging benchmarks, HateMM and MultiHateClip, demonstrate that LELA outperforms all existing training-free baselines by a large margin. We also provide extensive ablations and qualitative visualizations, establishing LELA as a strong foundation for scalable and interpretable hate video localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。