arXiv:2604.04467cs.CV2026-04中稿 · CVPR

无需标注,通过动态与群体特征自监督学习群体活动表示。

Group-DINOmics: Incorporating People Dynamics into DINO for Self-supervised Group Activity Feature Learning

  • 引入人流动态与群体相关物体位置作为预训练任务。
  • 在多个公开数据集上达到当前最优的群体活动检索与识别性能。
  • 适合对自监督视觉表征与群体行为理解感兴趣的研究者。

本文提出一种无需群体活动标注的群体活动特征(GAF)学习方法。与以往依赖低层静态局部特征的方法不同,本工作利用动态感知和群体感知的预训练任务,结合DINO提供的局部与全局特征,实现对群体动态敏感的GAF学习。预训练任务分别采用人物流估计(捕获个体局部运动)和群体相关物体位置估计(建模场景上下文,如人与物体的空间关系),以分别捕捉局部动态与全局群体特征。在多个公开数据集上的实验表明,该方法在群体活动检索与识别任务中达到当前最优性能。消融实验证明了各组件的有效性。代码已开源:https://github.com/tezuka0001/Group-DINOmics。

原文摘要 · Abstract (English)

This paper proposes Group Activity Feature (GAF) learning without group activity annotations. Unlike prior work, which uses low-level static local features to learn GAFs, we propose leveraging dynamics-aware and group-aware pretext tasks, along with local and global features provided by DINO, for group-dynamics-aware GAF learning. To adapt DINO and GAF learning to local dynamics and global group features, our pretext tasks use person flow estimation and group-relevant object location estimation, respectively. Person flow estimation is used to represent the local motion of each person, which is an important cue for understanding group activities. In contrast, group-relevant object location estimation encourages GAFs to learn scene context (e.g., spatial relations of people and objects) as global features. Comprehensive experiments on public datasets demonstrate the state-of-the-art performance of our method in group activity retrieval and recognition. Our ablation studies verify the effectiveness of each component in our method. Code: https://github.com/tezuka0001/Group-DINOmics.

自监督学习群体行为视觉表征DINO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。