arXiv:2604.17062cs.CV2026-04中稿 · ICASSP 2026

用运动分离与负提示提升零样本视频动作识别

Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition

论文配图:Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition
图 1 · 摘自论文原文
  • 分离运动与静态特征,避免信息冗余
  • 通过正负提示对齐语义,增强泛化能力
  • 在细粒度数据集上表现优异,适合跨类别识别

由于已见类与未见类之间存在语义鸿沟,零样本动作识别极具挑战。本文提出一种新框架,通过解耦嵌入与语义引导交互增强CLIP模型。运动分离模块(MSM)将运动敏感特征与全局静态特征分离,运动聚合块(MAB)采用门控交叉注意力优化运动表征,避免冗余信息重新耦合。为提升对未见类别的泛化能力,通过将视频特征投影嵌入与正向文本提示对齐,并利用负向提示显式建模“非类别”语义。在标准基准测试中,本方法持续优于现有基于CLIP的模型,在粗粒度与细粒度数据集上均实现稳健的零样本动作识别性能。

原文摘要 · Abstract (English)

Zero-shot action recognition is challenging due to the semantic gap between seen and unseen classes. We present a novel framework that enhances CLIP with disentangled embeddings and semantic-guided interaction. A Motion Separation Module (MSM) separates motion-sensitive and global-static features, while a Motion Aggregation Block (MAB) employs gated cross-attention to refine motion representation without re-coupling redundant information. To facilitate generalization to unseen categories, we enforce semantic alignment between video features and textual representations by aligning projected embeddings with positive textual prompts, while leveraging negative prompts to explicitly model "non-class" semantics. Experiments on standard benchmarks demonstrate that our method consistently outperforms prior CLIP-based approaches, achieving robust zero-shot action recognition across both coarse and fine-grained datasets.

零样本识别视频理解CLIP语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。