arXiv:2512.16504cs.CV2025-12中稿 · ICPR'26

通过多尺度特征融合与片段对比学习,提升骨骼动作定位精度。

Skeleton-Snippet Contrastive Learning with Multiscale Feature Fusion for Action Localization

  • 设计片段判别预训练任务,增强时序敏感特征
  • 在BABEL数据集上实现领先定位性能,跨子集泛化强
  • 适合需要高精度动作边界检测的研究者

自监督预训练范式在基于骨骼的动作识别中已取得显著成效,但针对骨骼序列的时间动作定位仍具挑战且研究不足。与视频级动作识别不同,定位需捕捉相邻帧间标签变化的细微差异,依赖时序敏感特征。为此,我们提出一种片段判别预训练任务:将骨骼序列密集划分为非重叠片段,通过对比学习促进跨视频的片段区分能力。同时,采用U型模块融合骨干网络中间特征,提升帧级定位的特征分辨率。该方法在BABEL数据集上多个子集和评估协议下均优于现有对比学习方法,并在使用NTU RGB+D与BABEL预训练后,于PKUMMD上实现最优迁移学习性能。

原文摘要 · Abstract (English)

The self-supervised pretraining paradigm has achieved great success in learning 3D action representations for skeleton-based action recognition using contrastive learning. However, learning effective representations for skeleton-based temporal action localization remains challenging and underexplored. Unlike video-level {action} recognition, detecting action boundaries requires temporally sensitive features that capture subtle differences between adjacent frames where labels change. To this end, we formulate a snippet discrimination pretext task for self-supervised pretraining, which densely projects skeleton sequences into non-overlapping segments and promotes features that distinguish them across videos via contrastive learning. Additionally, we build on strong backbones of skeleton-based action recognition models by fusing intermediate features with a U-shaped module to enhance feature resolution for frame-level localization. Our approach consistently improves existing skeleton-based contrastive learning methods for action localization on BABEL across diverse subsets and evaluation protocols. We also achieve state-of-the-art transfer learning performance on PKUMMD with pretraining on NTU RGB+D and BABEL.

动作定位对比学习骨骼序列多尺度融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。