arXiv:2508.04900cs.CVcs.AI2025-08被引 3

发现视频仇恨内容标注存在时间噪声,影响模型判断可靠性。

Revealing Temporal Label Noise in Multimodal Hateful Video Classification

  • 用标注时间戳裁剪仇恨视频,聚焦真实仇恨片段。
  • 发现粗粒度标注导致模型分类置信度下降,决策边界混乱。
  • 适合研究视频仇恨检测与多模态模型鲁棒性的学者参考。

在线多媒体内容的快速传播加剧了仇恨言论的扩散,带来重大社会与监管挑战。尽管近期研究在多模态仇恨视频检测方面取得进展,但多数方法依赖粗粒度的视频级标注,忽略了仇恨内容的时间粒度,导致显著的标签噪声——被标记为仇恨的视频常包含大量非仇恨片段。本文通过细粒度分析,利用HateMM和MultiHateClip数据集中的标注时间戳,裁剪出明确的仇恨片段,并对这些片段进行探索性分析,揭示了仇恨与非仇恨内容在语义上的重叠及粗粒度标注带来的混淆。受控实验表明,时间戳噪声从根本上改变了模型的决策边界,削弱了分类置信度,凸显了仇恨表达的高度情境依赖性和时间连续性。研究结果为多模态仇恨视频的时间动态提供了新见解,强调需构建具备时间感知能力的模型与基准以提升鲁棒性与可解释性。代码与数据见:https://github.com/Multimodal-Intelligence-Lab-MIL/HatefulVideoLabelNoise。

原文摘要 · Abstract (English)

The rapid proliferation of online multimedia content has intensified the spread of hate speech, presenting critical societal and regulatory challenges. While recent work has advanced multimodal hateful video detection, most approaches rely on coarse, video-level annotations that overlook the temporal granularity of hateful content. This introduces substantial label noise, as videos annotated as hateful often contain long non-hateful segments. In this paper, we investigate the impact of such label ambiguity through a fine-grained approach. Specifically, we trim hateful videos from the HateMM and MultiHateClip English datasets using annotated timestamps to isolate explicitly hateful segments. We then conduct an exploratory analysis of these trimmed segments to examine the distribution and characteristics of both hateful and non-hateful content. This analysis highlights the degree of semantic overlap and the confusion introduced by coarse, video-level annotations. Finally, controlled experiments demonstrated that time-stamp noise fundamentally alters model decision boundaries and weakens classification confidence, highlighting the inherent context dependency and temporal continuity of hate speech expression. Our findings provide new insights into the temporal dynamics of multimodal hateful videos and highlight the need for temporally aware models and benchmarks for improved robustness and interpretability. Code and data are available at https://github.com/Multimodal-Intelligence-Lab-MIL/HatefulVideoLabelNoise.

仇恨视频多模态时间噪声标注质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。