arXiv:2604.15196cs.CV2026-04

通过分层时空向量量化,无监督地分割骨架动作序列。

Unsupervised Skeleton-Based Action Segmentation via Hierarchical Spatiotemporal Vector Quantization

论文配图:Unsupervised Skeleton-Based Action Segmentation via Hierarchical Spatiotemporal Vector Quantization
图 1 · 摘自论文原文
  • 分两层量化:先分细粒度子动作,再聚合成动作级表征。
  • 在多个数据集上超越现有方法,减少片段长度偏差。
  • 适合无标注骨架动作分割任务,尤其关注时序结构建模。

我们提出一种新颖的分层时空向量量化框架,用于无监督的骨架动作时间分割。首先引入分层策略,包含两个连续的向量量化层级:低层将骨架关联到细粒度子动作,高层进一步将子动作聚合为动作级表示。该分层方法优于非分层基线,主要依赖空间线索重建输入骨架。随后,通过融合空间与时间信息,构建分层时空向量量化方案。该方法实现多层级聚类,同时恢复输入骨架及其对应的时间戳。在HuGaDB、LARa和BABEL等多个基准上的大量实验表明,该方法达到新最优性能,并有效降低无监督骨架动作分割中的片段长度偏差。

原文摘要 · Abstract (English)

We propose a novel hierarchical spatiotemporal vector quantization framework for unsupervised skeleton-based temporal action segmentation. We first introduce a hierarchical approach, which includes two consecutive levels of vector quantization. Specifically, the lower level associates skeletons with fine-grained subactions, while the higher level further aggregates subactions into action-level representations. Our hierarchical approach outperforms the non-hierarchical baseline, while primarily exploiting spatial cues by reconstructing input skeletons. Next, we extend our approach by leveraging both spatial and temporal information, yielding a hierarchical spatiotemporal vector quantization scheme. In particular, our hierarchical spatiotemporal approach performs multi-level clustering, while simultaneously recovering input skeletons and their corresponding timestamps. Lastly, extensive experiments on multiple benchmarks, including HuGaDB, LARa, and BABEL, demonstrate that our approach establishes a new state-of-the-art performance and reduces segment length bias in unsupervised skeleton-based temporal action segmentation.

动作分割骨架分析无监督学习向量量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。