通过分层向量量化提升动作片段分割的准确性与泛化能力
Hierarchical Vector Quantization for Unsupervised Action Segmentation
- 采用两级向量量化实现分层聚类,捕捉同类动作内部差异
- 在三个数据集上F1、召回率与JSD指标均优于现有方法
- 新引入的JSD度量更准确反映片段长度分布一致性
本文研究无监督时间动作分割任务,旨在将长时未修剪视频划分为语义一致的动作片段。现有方法虽融合表征学习与聚类,但难以处理同一类别内的时间片段变化。为此,提出分层向量量化(HVQ)方法,包含两个连续的向量量化模块,形成层级聚类结构,使子簇覆盖主簇内的变化。实验表明,该方法在片段长度分布建模上显著优于当前最优。为此,引入基于詹森-香农散度(JSD)的新评估指标。在Breakfast、YouTube Instructional和IKEA ASM三个公开数据集上验证,本方法在F1分数、召回率和JSD上均取得最佳表现。
原文摘要 · Abstract (English)
In this work, we address unsupervised temporal action segmentation, which segments a set of long, untrimmed videos into semantically meaningful segments that are consistent across videos. While recent approaches combine representation learning and clustering in a single step for this task, they do not cope with large variations within temporal segments of the same class. To address this limitation, we propose a novel method, termed Hierarchical Vector Quantization (HVQ), that consists of two subsequent vector quantization modules. This results in a hierarchical clustering where the additional subclusters cover the variations within a cluster. We demonstrate that our approach captures the distribution of segment lengths much better than the state of the art. To this end, we introduce a new metric based on the Jensen-Shannon Distance (JSD) for unsupervised temporal action segmentation. We evaluate our approach on three public datasets, namely Breakfast, YouTube Instructional and IKEA ASM. Our approach outperforms the state of the art in terms of F1 score, recall and JSD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。