让模型像人一样解释为何判定视频仇恨,提升可解释性。
Decoding Multimodal Cues: Unveiling the Implicit Meaning Behind Hateful Videos

- 用多模态思维链整合有害元素,增强推理证据
- 通过偏好优化使模型生成更符合逻辑的解释
- 新数据集+新框架,兼顾准确率与解释力
仇恨视频在在线平台日益泛滥,亟需有效检测。现有研究多聚焦二分类,缺乏对判断背后隐含意义的上下文解释,严重削弱模型可解释性。为此,本文提出可解释的仇恨视频检测方法,使模型不仅能作出判断,还能提供融合证据与逻辑推理的上下文理由。首先构建两个新数据集:Ex-HateMM 和 Ex-ImpliHateVid,包含细粒度的多模态有害元素标注及上下文理由。随后提出信息增强与推理优化(IARE)框架:第一阶段利用多模态思维链整合有害元素以丰富理由证据;第二阶段采用直接偏好优化,引导模型走向正确推理路径,避免错误逻辑,提升解释一致性。在两个数据集上进行大量实验,结果表明 IARE 在性能上达到当前最优,并能生成准确合理的解释。
原文摘要 · Abstract (English)
Hateful videos have become prevalent on online platforms, highlighting an urgent need for effective detection. However, existing studies primarily focus on binary classification and fail to provide contextual rationales that reveal the implicit meanings behind these judgments, significantly undermining model explainability. To fill this gap, we aim to achieve explainable hateful video detection, enabling models to provide contextual rationales that integrate relevant evidence and logical reasoning alongside decisions. This approach can comprehensively enhance the understanding of video content and the explainability of the decision-making process. We first introduce two datasets, Ex-HateMM and Ex-ImpliHateVid, for explainable hateful video detection. Each dataset provides fine-grained annotations of multimodal harmful elements, along with contextual rationales. We then propose an Information Augmentation and Reasoning Enhancement (IARE) framework designed for explainable detection. The framework employs an information augmentation phase that leverages the multimodal chain-of-thought to integrate harmful elements, thereby enriching rationale evidence. Additionally, IARE incorporates a reasoning enhancement phase, in which Direct Preference Optimization guides the model toward correct reasoning paths and away from incorrect ones, thereby improving the logical coherence of its justifications. We conduct extensive experiments on the two datasets, comparing multiple baselines with our proposed IARE framework. The results demonstrate that IARE achieves state-of-the-art performance while also generating accurate rationales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。