arXiv:2412.16946cs.CV2024-12被引 3

解决家庭环境视频动作识别中的持续学习难题

Video Domain Incremental Learning for Human Action Recognition in Home Environments

  • 提出视频领域增量学习框架,支持跨场景持续学习
  • 三种分割方式评估真实家庭环境下的领域漂移挑战
  • 仅用重放与采样策略即超越多数现有方法

由于非受限家庭环境中人类日常行为的多样性与动态变化,准确识别居家动作极具挑战性,亟需模型持续适应不同用户与场景。当前视频理解模型在新领域上微调常引发灾难性遗忘,导致旧场景性能下降。为此,本文正式提出视频领域增量学习(VDIL)问题,使模型能持续学习不同领域,同时保持固定的动作类别集合。现有持续学习研究多集中于类别增量学习,而视频理解中的领域增量学习尚未被充分关注。本文构建了一个面向非受限家庭环境的新型领域增量动作识别基准,设计了用户、场景、混合三种领域划分方式,系统评估真实场景下领域偏移带来的挑战。此外,提出一种无需领域标签的基线学习策略,结合重放与水库采样技术,在有限记忆和任务无关场景中表现优异。大量实验表明,该简单采样与重放策略在三个基准上均优于多数现有持续学习方法。

原文摘要 · Abstract (English)

It is significantly challenging to recognize daily human actions in homes due to the diversity and dynamic changes in unconstrained home environments. It spurs the need to continually adapt to various users and scenes. Fine-tuning current video understanding models on newly encountered domains often leads to catastrophic forgetting, where the models lose their ability to perform well on previously learned scenarios. To address this issue, we formalize the problem of Video Domain Incremental Learning (VDIL), which enables models to learn continually from different domains while maintaining a fixed set of action classes. Existing continual learning research primarily focuses on class-incremental learning, while the domain incremental learning has been largely overlooked in video understanding. In this work, we introduce a novel benchmark of domain incremental human action recognition for unconstrained home environments. We design three domain split types (user, scene, hybrid) to systematically assess the challenges posed by domain shifts in real-world home settings. Furthermore, we propose a baseline learning strategy based on replay and reservoir sampling techniques without domain labels to handle scenarios with limited memory and task agnosticism. Extensive experimental results demonstrate that our simple sampling and replay strategy outperforms most existing continual learning methods across the three proposed benchmarks.

动作识别持续学习视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。