通过构建时空相似性体积,提升开放词汇动作识别的细粒度匹配能力。
Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition

- 构建局部视频片段与动作类别的4维时空相似性体积
- 在多个基准上实现零样本、少样本任务的竞争力表现
- 适合需要细粒度动作理解的视频分析场景
当前开放词汇动作识别方法通常在计算文本对齐前将视觉特征全局聚合,这会丢失局部补丁信息和细粒度时空线索。本文提出相似性体积聚合(SimVA),从局部视觉-文本相似性构建密集的4维时空相似性体积。SimVA在局部视频标记与动作类别之间构建时空相似性体积,并采用类别采样确保大词表下的可扩展性。通过空间聚合对相似性体积进行精炼,增强帧内一致性;运动感知调制注入帧间变化线索,突出动态变化区域。最后使用Mamba进行时间聚合,建模跨帧的类别条件相似性演化。通过保持密集的视觉-文本对应关系,SimVA有效将CLIP迁移至视频动作识别,在零样本、少样本及基类到新类等多个基准上取得竞争力表现。
原文摘要 · Abstract (English)
Recent Open-Vocabulary Action Recognition (OVAR) methods typically aggregate visual features into a global representation before computing text alignment, a process that obscures local patch information and fine-grained spatio-temporal cues. We propose Similarity Volume Aggregation (SimVA), a framework that constructs a dense 4D spatio-temporal similarity volume from patch-level visual-text similarities. SimVA constructs a spatio-temporal similarity volume over local video tokens and action classes, and employs class sampling to ensure similarity aggregation scalable to large vocabularies. The similarity volume is refined by spatial aggregation, which contextualizes local similarity patterns to improve intra-frame consistency. Motion-aware modulation further injects inter-frame variation cues, highlighting dynamically changing regions. Mamba-based temporal aggregation then models the evolution of class-conditioned similarity patterns across frames. By maintaining dense visual-text correspondence, SimVA effectively transfers CLIP to video action recognition, achieving competitive performance across zero-shot, few-shot, and base-to-novel benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。