arXiv:2608.25701cs.CV2026-08

用弱监督预训练实现零样本动作定位,省去大量标注成本

Skeleton-based Zero-Shot Spatio-Temporal Action Localization via Weakly-Supervised Pretraining

论文配图:Skeleton-based Zero-Shot Spatio-Temporal Action Localization via Weakly-Supervised Pretraining
图 1 · 摘自论文原文
  • 通过视觉语言预训练,将骨架特征与动作文本对齐
  • 在四个数据集上实现零样本定位,准确率优于现有方法
  • 适合动作识别少样本或无标注场景的应用

我们提出一种基于骨架的零样本时空动作定位新预训练策略,通过大规模动作场景数据集进行弱监督预训练,以估计未见动作的人体实例,同时克服训练时高标注成本问题。具体而言,所提方法名为骨架-语言特征池化切换(Skeleton-Language feature Pooling Switching),引入弱监督视觉-语言预训练机制:预训练阶段,池化核在视频层面聚合骨架特征,并与视频已知动作文本嵌入对齐;推理阶段则无需训练,通过目标动作计算实例级特征。此外,我们提出场景混合判别对比学习,在多实例学习(MIL)框架下区分实例级动作。在四个公开的时空动作定位与分类数据集上的实验表明,该方法有效缓解标注瓶颈。

原文摘要 · Abstract (English)

We propose a novel pretraining strategy for skeleton-based zero-shot spatio-temporal action localization to estimate unseen actions for person instances while overcoming high annotation costs for training via new target actions and pretraining using large-scale action scenery datasets. Specifically, our approach, termed Skeleton-Language feature Pooling Switching, introduces a weakly-supervised vision-language pretraining mechanism. This mechanism transitions pooling kernels from pretraining, which aggregates skeleton features at the video level and aligns them with each video's known action text embeddings, to the inference phase that computes instance-level features without training via target actions. Furthermore, we propose Scene-Mixed Discriminative Contrastive Learning to distinguish actions at the instance level within the combined scene through the MIL framework. Our experiments on four public spatio-temporal action localization and classification datasets demonstrate that the proposed method effectively addresses annotation limitations.

动作定位零样本弱监督骨架模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。