arXiv:2604.20760cs.CV2026-04

通过多阶自相似性提升视频动作理解能力

Exploring High-Order Self-Similarity for Video Understanding

论文配图:Exploring High-Order Self-Similarity for Video Understanding
图 1 · 摘自论文原文
  • 设计轻量级MOSS模块,融合不同阶次的时空自相似特征
  • 在动作识别等任务中显著提升性能,计算开销极低
  • 适合需要高效建模运动动态的视频理解场景

时空自相似性(STSS)通过捕捉帧间的视觉对应关系,为视频理解提供有效的时序动态表征。本文探索更高阶的STSS,揭示其在不同阶次下展现的动态特性差异。提出多阶自相似性(MOSS)模块,一种轻量级神经模块,可学习并整合多阶STSS特征。该模块适用于多种视频任务,显著增强运动建模能力,同时仅带来微小的计算与内存开销。在视频动作识别、以运动为中心的视频问答及真实机器人任务上的大量实验均显示显著性能提升,验证了MOSS作为通用时序建模模块的广泛适用性。代码与模型权重将公开。

原文摘要 · Abstract (English)

Space-time self-similarity (STSS), which captures visual correspondences across frames, provides an effective way to represent temporal dynamics for video understanding. In this work, we explore higher-order STSS and demonstrate how STSSs at different orders reveal distinct aspects of these dynamics. We then introduce the Multi-Order Self-Similarity (MOSS) module, a lightweight neural module designed to learn and integrate multi-order STSS features. It can be applied to diverse video tasks to enhance motion modeling capabilities while consuming only marginal computational cost and memory usage. Extensive experiments on video action recognition, motion-centric video VQA, and real-world robotic tasks consistently demonstrate substantial improvements, validating the broad applicability of MOSS as a general temporal modeling module. The source code and checkpoints will be publicly available.

视频理解自相似性运动建模轻量模块

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。