arXiv:2512.02743cs.CVcs.AI2025-12中稿 · Transactions on Ma…被引 1

通过推理感知融合提升仇恨视频检测准确率

Reasoning-Aware Multimodal Fusion for Hateful Video Detection

  • 设计局部-全局上下文融合与语义交叉注意力,增强多模态交互
  • 在两个真实数据集上,宏平均F1提升3%,仇恨类召回率提升7%
  • 适合关注多模态内容理解与仇恨识别的工程师和研究者

在线视频中的仇恨言论正对数字平台构成日益严重的威胁,尤其随着视频内容日益多模态和依赖上下文。现有方法往往难以有效融合不同模态间的复杂语义关系,且缺乏对细微仇恨内容的理解能力。为此,我们提出一种新型推理感知多模态融合(RAMF)框架。为解决第一问题,设计局部-全局上下文融合(LGCF)以捕捉局部显著线索与全局时序结构,并提出语义交叉注意力(SCA)实现细粒度多模态语义交互。为应对第二挑战,引入对抗性推理机制——一种三阶段结构化过程:视觉语言模型生成(i)客观描述,(ii)假设仇恨的推断,(iii)非仇恨假设的推断,提供互补的语义视角,丰富模型对细微仇恨意图的上下文理解。在两个真实世界仇恨视频数据集上的评估表明,本方法实现稳健泛化性能,在宏平均F1和仇恨类召回率上分别优于最先进方法3%和7%。源代码及复现所需数据见https://github.com/Multimodal-Intelligence-Lab-MIL/RAMF。

原文摘要 · Abstract (English)

Hate speech in online videos is posing an increasingly serious threat to digital platforms, especially as video content becomes increasingly multimodal and context-dependent. Existing methods often struggle to effectively fuse the complex semantic relationships between modalities and lack the ability to understand nuanced hateful content. To address these issues, we propose an innovative Reasoning-Aware Multimodal Fusion (RAMF) framework. To tackle the first challenge, we design Local-Global Context Fusion (LGCF) to capture both local salient cues and global temporal structures, and propose Semantic Cross Attention (SCA) to enable fine-grained multimodal semantic interaction. To tackle the second challenge, we introduce adversarial reasoning-a structured three-stage process where a vision-language model generates (i) objective descriptions, (ii) hate-assumed inferences, and (iii) non-hate-assumed inferences-providing complementary semantic perspectives that enrich the model's contextual understanding of nuanced hateful intent. Evaluations on two real-world hateful video datasets demonstrate that our method achieves robust generalisation performance, improving upon state-of-the-art methods by 3% and 7% in Macro-F1 and hate class recall, respectively. The source codes and data required to reproduce our results are available at https://github.com/Multimodal-Intelligence-Lab-MIL/RAMF.

多模态融合仇恨检测视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。