arXiv:2503.16832cs.CV2025-03ICCV被引 13

用统一模型同时实现视频对齐与动作分割,效率更高。

Joint Self-Supervised Video Alignment and Action Segmentation

  • 基于融合结构先验的最优传输框架,高效训练且迭代少。
  • 单任务对齐性能达新高,多任务下动作分割更优。
  • 首个将对齐与分割统一建模的方法,节省时间和内存。

我们提出一种基于统一最优传输框架的新型自监督视频对齐与动作分割方法。首先,通过引入融合结构先验的混合Gromov-Wasserstein最优传输公式,实现了高效的自监督视频对齐训练,仅需少量迭代即可求解,且在GPU上运行迅速。该单任务方法在多个视频对齐基准上达到当前最优性能,优于依赖传统Kantorovich最优传输与最优性先验的VAVA方法。进一步地,我们提出统一的最优传输框架,实现视频对齐与动作分割的联合建模,仅需训练和存储一个模型,相比两个独立模型显著降低时间与内存开销。在多个视频对齐与动作分割数据集上的大量实验表明,该多任务方法在对齐性能相当的前提下,动作分割效果显著优于以往方法。据我们所知,这是首个将视频对齐与动作分割统一为单一模型的工作。代码已公开于https://retrocausal.ai/research/。

原文摘要 · Abstract (English)

We introduce a novel approach for simultaneous self-supervised video alignment and action segmentation based on a unified optimal transport framework. In particular, we first tackle self-supervised video alignment by developing a fused Gromov-Wasserstein optimal transport formulation with a structural prior, which trains efficiently on GPUs and needs only a few iterations for solving the optimal transport problem. Our single-task method achieves the state-of-the-art performance on multiple video alignment benchmarks and outperforms VAVA, which relies on a traditional Kantorovich optimal transport formulation with an optimality prior. Furthermore, we extend our approach by proposing a unified optimal transport framework for joint self-supervised video alignment and action segmentation, which requires training and storing a single model and saves both time and memory consumption as compared to two different single-task models. Extensive evaluations on several video alignment and action segmentation datasets demonstrate that our multi-task method achieves comparable video alignment yet superior action segmentation results over previous methods in video alignment and action segmentation respectively. Finally, to the best of our knowledge, this is the first work to unify video alignment and action segmentation into a single model. Our code is available on our research website: https://retrocausal.ai/research/.

视频对齐动作分割最优传输自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。