arXiv:2508.15389cs.CV2025-08中稿 · IEEE TIP被引 13

用脉冲神经网络提升视频摘要的关键词帧提取与语义连贯性

Spiking Variational Graph Representation Inference for Video Summarization

  • 基于脉冲神经网络设计事件驱动的关键帧提取器
  • 动态图推理模块解耦物体一致性与语义连贯性,提升细粒度推理
  • 变分推断重建模块降低多通道特征融合中的噪声影响

随着短视频内容兴起,高效提取关键信息的视频摘要技术愈发重要。现有方法难以捕捉全局时间依赖性,且在多通道特征融合中易受噪声干扰,影响语义连贯性。本文提出脉冲变分图网络(SpiVG),通过脉冲神经网络(SNN)的事件驱动机制实现自主学习关键帧特征;引入动态聚合图推理器,解耦上下文物体一致性与语义视角连贯性,支持细粒度自适应推理;设计变分推断重建模块,利用证据下界优化(ELBO)捕获多通道特征分布的潜在结构,并通过后验分布正则化抑制过拟合。实验表明,SpiVG在SumMe、TVSum、VideoXum、QFVS等多个数据集上优于现有方法。代码与预训练模型已开源。

原文摘要 · Abstract (English)

With the rise of short video content, efficient video summarization techniques for extracting key information have become crucial. However, existing methods struggle to capture the global temporal dependencies and maintain the semantic coherence of video content. Additionally, these methods are also influenced by noise during multi-channel feature fusion. We propose a Spiking Variational Graph (SpiVG) Network, which enhances information density and reduces computational complexity. First, we design a keyframe extractor based on Spiking Neural Networks (SNN), leveraging the event-driven computation mechanism of SNNs to learn keyframe features autonomously. To enable fine-grained and adaptable reasoning across video frames, we introduce a Dynamic Aggregation Graph Reasoner, which decouples contextual object consistency from semantic perspective coherence. We present a Variational Inference Reconstruction Module to address uncertainty and noise arising during multi-channel feature fusion. In this module, we employ Evidence Lower Bound Optimization (ELBO) to capture the latent structure of multi-channel feature distributions, using posterior distribution regularization to reduce overfitting. Experimental results show that SpiVG surpasses existing methods across multiple datasets such as SumMe, TVSum, VideoXum, and QFVS. Our codes and pre-trained models are available at https://github.com/liwrui/SpiVG.

视频摘要脉冲神经网络图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。