用双空间几何学习提升弱监督视频暴力检测的区分能力
Beyond Euclidean: Dual-Space Representation Learning for Weakly Supervised Video Violence Detection
- 在欧氏与双曲空间并行学习视觉特征与事件关系
- 通过双曲注意力机制提升相似场景的区分度
- 适合需要精准识别模糊暴力行为的研究者
尽管众多视频暴力检测(VVD)方法聚焦于欧氏空间中的表示学习,但其难以学习到足够区分性的特征,导致在识别与暴力行为视觉相似的正常事件时表现不佳(即模糊暴力)。相比之下,双曲表示学习因其能建模事件间的层次与复杂关系,具备增强相似事件区分能力的潜力。受此启发,本文提出一种新型双空间表示学习(DSRL)方法,用于弱监督视频暴力检测,结合欧氏与双曲几何的优势,既捕捉事件的视觉特征,又探索事件间的内在关系,从而提升特征的区分能力。DSRL采用新颖的信息聚合策略,在双曲空间中逐层学习事件上下文,通过层敏感的双曲关联度选择聚合节点,并受双曲Dirichlet能量约束。此外,为打破空间间的信息隔阂,引入跨空间注意力机制,促进欧氏与双曲空间间的交互,以获取更优的判别性特征。大量实验验证了所提方法的有效性。
原文摘要 · Abstract (English)
While numerous Video Violence Detection (VVD) methods have focused on representation learning in Euclidean space, they struggle to learn sufficiently discriminative features, leading to weaknesses in recognizing normal events that are visually similar to violent events (\emph{i.e.}, ambiguous violence). In contrast, hyperbolic representation learning, renowned for its ability to model hierarchical and complex relationships between events, has the potential to amplify the discrimination between visually similar events. Inspired by these, we develop a novel Dual-Space Representation Learning (DSRL) method for weakly supervised VVD to utilize the strength of both Euclidean and hyperbolic geometries, capturing the visual features of events while also exploring the intrinsic relations between events, thereby enhancing the discriminative capacity of the features. DSRL employs a novel information aggregation strategy to progressively learn event context in hyperbolic spaces, which selects aggregation nodes through layer-sensitive hyperbolic association degrees constrained by hyperbolic Dirichlet energy. Furthermore, DSRL attempts to break the cyber-balkanization of different spaces, utilizing cross-space attention to facilitate information interactions between Euclidean and hyperbolic space to capture better discriminative features for final violence detection. Comprehensive experiments demonstrate the effectiveness of our proposed DSRL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。