仅用每段动作一个标注帧,实现高精度人体动作分割
Point-Supervised Skeleton-Based Human Action Segmentation
- 用单帧标签+多模态骨骼信息生成伪标签
- 在PKU-MMD等数据集上超越部分全监督方法
- 适合标注成本敏感的动作识别场景
基于骨骼的时间动作分割是智能系统感知人类行为的关键任务。全监督方法虽性能良好,但需昂贵的帧级标注且对模糊动作边界敏感。为此,本文提出一种点监督框架,仅需每动作段一个标注帧。利用预训练统一模型编码关节、骨骼和运动多模态骨骼数据,提取丰富特征。提出新型原型相似性方法,并与能量函数、约束K-Medoids聚类结合生成可靠伪标签。引入多模态伪标签融合策略提升伪标签可靠性并指导模型训练。在PKU-MMD(X-Sub、X-View)、MCFS-22和MCFS-130上建立新基准,实现对比基线。大量实验表明,该方法性能优异,甚至超过部分全监督方法,同时显著降低标注成本。
原文摘要 · Abstract (English)
Skeleton-based temporal action segmentation is a fundamental yet challenging task, playing a crucial role in enabling intelligent systems to perceive and respond to human activities. While fully-supervised methods achieve satisfactory performance, they require costly frame-level annotations and are sensitive to ambiguous action boundaries. To address these issues, we introduce a point-supervised framework for skeleton-based action segmentation, where only a single frame per action segment is labeled. We leverage multimodal skeleton data, including joint, bone, and motion information, encoded via a pretrained unified model to extract rich feature representations. To generate reliable pseudo-labels, we propose a novel prototype similarity method and integrate it with two existing methods: energy function and constrained K-Medoids clustering. Multimodal pseudo-label integration is proposed to enhance the reliability of the pseudo-label and guide the model training. We establish new benchmarks on PKU-MMD (X-Sub and X-View), MCFS-22, and MCFS-130, and implement baselines for point-supervised skeleton-based human action segmentation. Extensive experiments show that our method achieves competitive performance, even surpassing some fully-supervised methods while significantly reducing annotation effort.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。