构建多层视频危害理解基准,评估大模型深层推理能力
HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models

- 设计三层评测体系:可见证据、片段内含义、片段外推理
- 19个主流模型平均准确率仅61.7%,展现深层理解短板
- 提出动态上下文检索方法BCR,性能提升至84.4%新高
大型视觉语言模型(LVLMs)在自动化内容审核中展现出巨大潜力,但现有研究存在两大缺陷:一是忽视有害视频的多层次特征,多以二分类任务为主,难以捕捉隐性或深层语境危害;二是缺乏解释性,仅评估是否正确标记,未说明原因,使模型可通过表面捷径得分。为此,我们提出HarmVideoBench,一个包含1,379个视频与4,137道多选题的多层诊断基准,涵盖可观测证据、片段内含义和片段外推理三个层级,旨在评估模型对有害视频的深层理解能力。我们对19个领先模型进行测评,发现其平均宏观准确率为61.7%。同时引入BCR方法,可预测推理边界并按需动态检索上下文,显著提升模型表现,使宏观平均准确率从61.7%提升至84.4%的当前最优水平。
原文摘要 · Abstract (English)
Large vision-language models (LVLMs) have recently shown immense potential in automated content moderation, sparking growing interest in developing harmful-video benchmarks. However, we identify two primary limitations in existing works: 1) The multi-layered characteristics of harmful videos are overlooked. Existing benchmarks predominantly formulate evaluation as a binary classification task, failing to capture implicit or deep contextual harms. 2) Explanatory rationales are completely absent. Current frameworks measure exclusively whether a model flags a video correctly rather than explaining why, turning evaluation into a black box where models can succeed through superficial shortcuts. To address these problems, we present HarmVideoBench, a multi-layered diagnostic benchmark comprising 1,379 videos paired with 4,137 multiple-choice questions. HarmVideoBench benchmarks three hierarchical dimensions: Observable Evidence, Clip-Internal Meaning, and Beyond-Clip Reasoning, aiming to evaluate models' deep understanding beyond surface cues with carefully balanced and curated samples. We evaluate 19 leading models on HarmVideoBench to assess their multidimensional understanding of harmful videos. Moreover, we introduce BCR, a benchmark-aligned method that predicts reasoning boundaries and dynamically retrieves context only when needed. Experimental results show that BCR substantially improves the base model's performance in harmful video understanding, raising the macro average from 61.7 percent to a state-of-the-art 84.4 percent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。