arXiv:2608.27879cs.CVcs.LG2026-08

发现视频暴力检测的评估指标被前置画面的源线索干扰,掩盖了真实表现差异。

What Do Interaction Representations Actually Measure? Pre-Event Separability in Weakly-Supervised Violence Detection

论文配图:What Do Interaction Representations Actually Measure? Pre-Event Separability in Weakly-Supervised Violence Detection
图 1 · 摘自论文原文
  • 用固定流程对比五种交互表示法,发现细粒度姿态不优于粗略框体几何
  • 前置帧中91%的判别力来自片头字幕、平台水印等源信息而非实际暴力事件
  • 建议直接检查前置窗口以诊断评估偏差,无需额外标注

人体姿态提供比粗略空间关系更详细的体态信息,但在下游流程固定时,这种细节是否带来更强区分能力尚不明确。我们通过早期暴力检测进行检验:在追踪器、时序头、监督方式、数据划分和评估均固定的前提下,比较五种交互表示——从粗略边界框几何到匹配容量的原始关节学习编码器。在视频级评估下,无一姿态表示超越粗略几何;尽管存在15个异常视频,仍无法排除微小效应。扩展至冻结视觉编码器后,在含137个异常视频的XD-Violence上,人物裁剪外观与整帧上下文均显著优于几何特征;而在UCF-Crime上两者相当,整帧上下文更优。进一步分析发现,仅使用标注起始前帧进行评分,去除序列长度提示后,仍保持39%-91%的超越随机水平的分离能力,包括七个手工设计的几何通道。紧致前置窗口分析揭示具体来源伪影:片头字幕与平台水印,这些在正常类监控视频中不存在。因此,视频级AUC实质是事件证据与前置源线索的混合信号,这种共享判别源会掩盖表示方法间的真正差异。诊断仅需利用现有基准自带标注。

原文摘要 · Abstract (English)

Articulated human pose provides detailed body-configuration information beyond coarse spatial relationships, but whether this detail yields greater discriminative information when the downstream pipeline is held fixed remains unclear. We examine this through early violence detection. Holding the tracker, temporal head, supervision, folds, and evaluation fixed, we compare five interaction representations spanning coarse bounding-box geometry, a matched handcrafted pose analogue, enriched pose descriptors, and a matched-capacity encoder learned from raw joints, under video-level evaluation with cluster-bootstrap intervals. No pose-based representation outperforms coarse geometry, though with fifteen anomalous videos this subset cannot rule out small effects. Extending the pipeline to frozen visual encoders, and repeating the comparison on XD-Violence (137 anomalous videos, nine times our UCF-Crime sample), person-crop appearance and whole-frame context both exceed geometry by a wide margin, yet context matches appearance on UCF-Crime and exceeds it on the larger split: cropping to the interacting people yields no advantage over encoding the whole frame. This prompts a direct test of what the benchmark measures. Scoring anomalous videos using only frames preceding the annotated onset, under a control removing sequence length as a cue, retains 39-91% of above-chance separation on both benchmarks, including for seven hand-designed geometric channels. Inspection of the tightest pre-onset windows identifies concrete provenance artifacts: editorial title cards and platform watermarks absent from the surveillance footage supplying the normal class. Video-level AUC here is thus a composite of event evidence and pre-event source cues, a shared source of discrimination that can obscure differences between representations. The diagnostic requires only annotations these benchmarks already ship.

暴力检测评估偏差源线索行为识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。