arXiv:2508.21773cs.CVcs.AI2025-08中稿 · The 36th British M…

无标签视频持续学习新方法,自动识别新任务并扩展记忆

Unsupervised Video Continual Learning via Non-Parametric Deep Embedded Clustering

  • 用深度嵌入特征的核密度估计构建非参数概率表示
  • 无需标签和任务边界,连续学习多个视频任务性能显著提升
  • 适合无监督视频理解、持续学习研究者参考

我们提出一种真实场景下的无监督视频持续学习框架,即在学习一系列任务时既无任务边界也无标签。针对视频数据复杂的时空特性及现有研究多依赖标签与任务边界的问题,我们首次探索无监督视频持续学习(uVCL)。视频处理带来更高计算与内存需求,为此我们设计了通用基准实验协议,考虑每个任务中无结构的视频类别学习。采用无监督视频变换网络提取深度嵌入特征,利用核密度估计(KDE)作为非参数概率表征,并引入新颖性检测准则,动态扩展记忆聚类以捕捉新知识。通过迁移前序任务的预训练状态实现知识传递。在UCF101、HMDB51、Something-to-Something V2三个标准视频动作识别数据集上进行评估,完全不使用标签或类别边界,结果表明该方法显著提升了连续学习多个任务的性能。

原文摘要 · Abstract (English)

We propose a realistic scenario for the unsupervised video learning where neither task boundaries nor labels are provided when learning a succession of tasks. We also provide a non-parametric learning solution for the under-explored problem of unsupervised video continual learning. Videos represent a complex and rich spatio-temporal media information, widely used in many applications, but which have not been sufficiently explored in unsupervised continual learning. Prior studies have only focused on supervised continual learning, relying on the knowledge of labels and task boundaries, while having labeled data is costly and not practical. To address this gap, we study the unsupervised video continual learning (uVCL). uVCL raises more challenges due to the additional computational and memory requirements of processing videos when compared to images. We introduce a general benchmark experimental protocol for uVCL by considering the learning of unstructured video data categories during each task. We propose to use the Kernel Density Estimation (KDE) of deep embedded video features extracted by unsupervised video transformer networks as a non-parametric probabilistic representation of the data. We introduce a novelty detection criterion for the incoming new task data, dynamically enabling the expansion of memory clusters, aiming to capture new knowledge when learning a succession of tasks. We leverage the use of transfer learning from the previous tasks as an initial state for the knowledge transfer to the current learning task. We found that the proposed methodology substantially enhances the performance of the model when successively learning many tasks. We perform in-depth evaluations on three standard video action recognition datasets, including UCF101, HMDB51, and Something-to-Something V2, without using any labels or class boundaries.

视频持续学习无监督学习聚类KDE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。