arXiv:2504.18866cs.CV2025-04TPAMI被引 11

用双空间几何增强视频暴力检测,特别擅长区分易混淆场景

PiercingEye: Dual-Space Video Violence Detection with Hyperbolic Vision-Language Guidance

  • 融合欧氏与双曲几何建模事件层级关系
  • 在两个数据集上达到当前最佳性能,尤其在模糊样本上提升显著
  • 适合需要细粒度识别暴力行为的安防与内容审核场景

现有弱监督视频暴力检测方法主要依赖欧氏表示学习,难以区分视觉相似但语义不同的事件,原因在于层次建模能力有限且模糊样本不足。为此,我们提出PiercingEye,一种结合欧氏与双曲几何的双空间学习框架,以增强特征判别力。具体而言,PiercingEye引入分层敏感的双曲聚合策略及双曲狄利克雷能量约束,逐步建模事件层级;同时设计跨空间注意力机制,促进欧氏与双曲空间间的互补特征交互。为缓解模糊样本稀缺问题,利用大语言模型生成逻辑引导的模糊事件描述,通过双曲视觉-语言对比损失实现显式监督,并采用动态相似性感知加权机制强化高混淆样本的学习。在XD-Violence和UCF-Crime基准上的大量实验表明,PiercingEye取得当前最优性能,尤其在新构建的模糊事件子集上表现突出,验证了其在细粒度暴力检测中的卓越能力。

原文摘要 · Abstract (English)

Existing weakly supervised video violence detection (VVD) methods primarily rely on Euclidean representation learning, which often struggles to distinguish visually similar yet semantically distinct events due to limited hierarchical modeling and insufficient ambiguous training samples. To address this challenge, we propose PiercingEye, a novel dual-space learning framework that synergizes Euclidean and hyperbolic geometries to enhance discriminative feature representation. Specifically, PiercingEye introduces a layer-sensitive hyperbolic aggregation strategy with hyperbolic Dirichlet energy constraints to progressively model event hierarchies, and a cross-space attention mechanism to facilitate complementary feature interactions between Euclidean and hyperbolic spaces. Furthermore, to mitigate the scarcity of ambiguous samples, we leverage large language models to generate logic-guided ambiguous event descriptions, enabling explicit supervision through a hyperbolic vision-language contrastive loss that prioritizes high-confusion samples via dynamic similarity-aware weighting. Extensive experiments on XD-Violence and UCF-Crime benchmarks demonstrate that PiercingEye achieves state-of-the-art performance, with particularly strong results on a newly curated ambiguous event subset, validating its superior capability in fine-grained violence detection.

视频暴力检测双曲几何弱监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。