通过正则化最优传输学习视频关键步骤与顺序
Procedure Learning via Regularized Gromov-Wasserstein Optimal Transport
- 用融合结构先验的格罗莫夫-沃瑟斯坦运输做帧间对齐
- 对比正则化防止所有帧挤成一团,避免退化解
- 在第一人称和第三人称数据集上效果优于已有方法
我们研究自监督流程学习,从无标签视频中发现关键步骤及其顺序。以往方法通常在确定关键步骤前学习视频间的帧对齐,但易受顺序变化、背景/冗余帧和重复动作影响。为此,我们提出一种自监督框架,采用融合结构先验的格罗莫夫-沃瑟斯坦最优传输进行帧对齐。然而,仅优化时间对齐可能导致退化解:所有帧映射到嵌入空间中一个极小聚类,使每段视频仅对应一个关键步骤。为解决此问题,我们引入对比正则化,促使不同帧映射到不同位置,避免平凡解。在第一人称和第三人称基准上的大量实验表明,本方法显著优于先前工作,包括依赖经典柯朗托维奇最优传输与最优性先验的OPEL。
原文摘要 · Abstract (English)
We study self-supervised procedure learning, which discovers key steps and their order from a set of unlabeled videos. Previous methods typically learn frame-to-frame correspondences between videos before determining key steps and their order. However, their performance often suffers from order variations, background/redundant frames, and repeated actions. To overcome these challenges, we propose a self-supervised framework, which utilizes a fused Gromov-Wasserstein optimal transport with a structural prior for frame-to-frame mapping. However, optimizing only for the above temporal alignment may lead to degenerate solutions, where all frames are mapped to a small cluster in the embedding space and thus every video is assigned to just one key step. To address that issue, we integrate a contrastive regularization, which maps different frames to various points, avoiding trivial solutions. Finally, extensive experiments on egocentric and third-person benchmarks demonstrate our superior performance over prior works, including OPEL which relies on a classical Kantorovich optimal transport with an optimality prior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。