提出闭环最优传输框架,提升无监督动作分割的时序一致性
CLOT: Closed Loop Optimal Transport for Unsupervised Action Segmentation
- 通过双层最优传输学习帧与片段嵌入及伪标签
- 引入循环机制,使帧嵌入与伪标签相互优化
- 在4个数据集上验证性能,适合时序动作分析研究
无监督动作分割近年取得进展,其中基于最优传输(OT)的方法ASOT能同时学习动作表征并进行聚类,无需假设动作顺序,可从帧与动作标签间的噪声代价矩阵中解码出时序一致的分割结果。然而,该方法缺乏片段级监督,导致帧与动作表征间反馈受限。为此,本文提出闭环最优传输(CLOT),一种基于OT的新框架,具备多层级循环特征学习机制。利用编码器-解码器结构,CLOT通过求解两个独立的OT问题同时学习伪标签、帧嵌入和片段嵌入,并通过帧与片段嵌入间的交叉注意力,结合第三个OT问题对两者进行精炼。在四个基准数据集上的实验表明,循环学习机制显著提升了无监督动作分割性能。
原文摘要 · Abstract (English)
Unsupervised action segmentation has recently pushed its limits with ASOT, an optimal transport (OT)-based method that simultaneously learns action representations and performs clustering using pseudo-labels. Unlike other OT-based approaches, ASOT makes no assumptions about action ordering and can decode a temporally consistent segmentation from a noisy cost matrix between video frames and action labels. However, the resulting segmentation lacks segment-level supervision, limiting the effectiveness of feedback between frames and action representations. To address this limitation, we propose Closed Loop Optimal Transport (CLOT), a novel OT-based framework with a multi-level cyclic feature learning mechanism. Leveraging its encoder-decoder architecture, CLOT learns pseudo-labels alongside frame and segment embeddings by solving two separate OT problems. It then refines both frame embeddings and pseudo-labels through cross-attention between the learned frame and segment embeddings, by integrating a third OT problem. Experimental results on four benchmark datasets demonstrate the benefits of cyclical learning for unsupervised action segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。