arXiv:2602.18043cs.CV2026-02TPAMI被引 16

用大模型拆解动作的时空属性,提升少样本动作识别效果

Spatio-temporal Decoupled Knowledge Compensator for Few-Shot Action Recognition

论文配图:Spatio-temporal Decoupled Knowledge Compensator for Few-Shot Action Recognition
图 1 · 摘自论文原文
  • 将动作名称分解为时空属性知识,增强语义理解
  • 通过时空补偿器学习物体级和帧级原型,捕捉细粒度特征
  • 在5个数据集上达到当前最优,适合少样本视觉识别研究者

少样本动作识别(FSAR)需仅用少量标注视频识别新动作类别。现有方法多以动作名称作为粗粒度上下文引导视觉特征学习,但此类信息不足以支持对动作中空间与时间概念的充分建模。本文提出DiST框架,通过分解-融合机制利用大语言模型提供的解耦时空知识,学习多粒度原型。在分解阶段,将原始动作名称拆分为多样化的时空属性描述(动作相关常识知识),从空间与时间双视角补充语义上下文。在融合阶段,设计空间知识补偿器(SKC)与时间知识补偿器(TKC),分别用于发现物体级和帧级原型:SKC基于空间知识自适应聚合关键图块令牌;TKC则利用时间属性辅助建模帧间时序关系。所学原型能有效揭示精细的空间细节与多样的时间模式。实验表明,DiST在五个标准FSAR数据集上均取得当前最优性能。

原文摘要 · Abstract (English)

Few-Shot Action Recognition (FSAR) is a challenging task that requires recognizing novel action categories with a few labeled videos. Recent works typically apply semantically coarse category names as auxiliary contexts to guide the learning of discriminative visual features. However, such context provided by the action names is too limited to provide sufficient background knowledge for capturing novel spatial and temporal concepts in actions. In this paper, we propose DiST, an innovative Decomposition-incorporation framework for FSAR that makes use of decoupled Spatial and Temporal knowledge provided by large language models to learn expressive multi-granularity prototypes. In the decomposition stage, we decouple vanilla action names into diverse spatio-temporal attribute descriptions (action-related knowledge). Such commonsense knowledge complements semantic contexts from spatial and temporal perspectives. In the incorporation stage, we propose Spatial/Temporal Knowledge Compensators (SKC/TKC) to discover discriminative object-level and frame-level prototypes, respectively. In SKC, object-level prototypes adaptively aggregate important patch tokens under the guidance of spatial knowledge. Moreover, in TKC, frame-level prototypes utilize temporal attributes to assist in inter-frame temporal relation modeling. These learned prototypes thus provide transparency in capturing fine-grained spatial details and diverse temporal patterns. Experimental results show DiST achieves state-of-the-art results on five standard FSAR datasets.

少样本识别时空建模大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。