arXiv:2510.12160cs.CV2025-10NeurIPS被引 2

通过时空信息聚合与传播,提升状态空间模型的视频理解能力。

State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video Understanding

  • 设计帧内聚合与帧间扩散模块,增强时空信息捕获
  • 在4个数据集上平均性能提升2.76%,参数量大幅减少
  • 适合需要高效微调的视频理解下游任务

近期预训练状态空间模型在视频分类中展现出巨大潜力,通过线性复杂度顺序压缩视频视觉标记,提升处理效率并保持高性能。为将强大预训练模型应用于下游任务,提示学习被提出以实现仅需少量可微调参数的高效适配。然而,顺序压缩的视觉提示标记难以捕捉视频中的空间与时间上下文信息,限制了帧内空间信息的有效传播和帧间时间信息的提取。为此,本文提出状态空间提示(SSP)方法,结合帧内与帧间提示,聚合并传播视频中的关键时空信息。具体而言,设计帧内聚集(IFG)模块以聚合每帧内的空间关键信息;同时设计帧间扩散(IFS)模块,将判别性时空信息跨帧传播。通过自适应平衡与压缩帧内及帧间的关键时空信息,所提SSP以互补方式有效传播视频中的判别性信息。大量实验在四个视频基准数据集上验证,其平均性能显著优于现有最先进方法2.76%,同时大幅降低微调参数开销。

原文摘要 · Abstract (English)

Recently, pre-trained state space models have shown great potential for video classification, which sequentially compresses visual tokens in videos with linear complexity, thereby improving the processing efficiency of video data while maintaining high performance. To apply powerful pre-trained models to downstream tasks, prompt learning is proposed to achieve efficient downstream task adaptation with only a small number of fine-tuned parameters. However, the sequentially compressed visual prompt tokens fail to capture the spatial and temporal contextual information in the video, thus limiting the effective propagation of spatial information within a video frame and temporal information between frames in the state compression model and the extraction of discriminative information. To tackle the above issue, we proposed a State Space Prompting (SSP) method for video understanding, which combines intra-frame and inter-frame prompts to aggregate and propagate key spatiotemporal information in the video. Specifically, an Intra-Frame Gathering (IFG) module is designed to aggregate spatial key information within each frame. Besides, an Inter-Frame Spreading (IFS) module is designed to spread discriminative spatio-temporal information across different frames. By adaptively balancing and compressing key spatio-temporal information within and between frames, our SSP effectively propagates discriminative information in videos in a complementary manner. Extensive experiments on four video benchmark datasets verify that our SSP significantly outperforms existing SOTA methods by 2.76% on average while reducing the overhead of fine-tuning parameters.

视频理解状态空间模型提示学习时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。