arXiv:2503.12519cs.CV2025-03

通过隐式聚类实现多动作序列对齐,提升模型泛化能力。

Multi Activity Sequence Alignment via Implicit Clustering

  • 隐式帧级聚类与双增强技术联合优化序列对齐
  • 在三个数据集上超越现有方法,支持多动作和多模态
  • 适合需要跨动作泛化的视频理解任务

自监督时序序列对齐可为多种应用提供丰富有效的表示。然而,现有方法大多仅限于同一动作序列的对齐,且需为每种动作单独训练模型。本文提出一种基于隐式聚类的序列对齐新框架,核心思想是在对齐序列帧的同时进行隐式片段级聚类。结合所提出的双增强技术,显著提升了网络学习通用且判别性表示的能力。实验表明,该方法在H2O、PennAction和IKEA ASM三个多样化数据集上均优于当前最优结果,充分展现了框架在多动作及不同模态下的泛化能力。代码将在论文录用后公开。

原文摘要 · Abstract (English)

Self-supervised temporal sequence alignment can provide rich and effective representations for a wide range of applications. However, existing methods for achieving optimal performance are mostly limited to aligning sequences of the same activity only and require separate models to be trained for each activity. We propose a novel framework that overcomes these limitations using sequence alignment via implicit clustering. Specifically, our key idea is to perform implicit clip-level clustering while aligning frames in sequences. This coupled with our proposed dual augmentation technique enhances the network's ability to learn generalizable and discriminative representations. Our experiments show that our proposed method outperforms state-of-the-art results and highlight the generalization capability of our framework with multi activity and different modalities on three diverse datasets, H2O, PennAction, and IKEA ASM. We will release our code upon acceptance.

序列对齐自监督多动作隐式聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。