arXiv:2410.23092cs.CV2024-10

针对道路场景原子动作识别,提出多分支框架与增强策略,获ECCV挑战赛第一名。

First Place Solution to the ECCV 2024 ROAD++ Challenge @ ROAD++ Atomic Activity Recognition 2024

  • 构建多分支框架,分离物体类别与单体/群体识别任务
  • 采用多策略集成,实现69% mAP的测试性能
  • 通过翻转视频帧与道路拓扑实现数据增广,缓解过拟合

本文介绍我们团队参与2024年ECCV ROAD++挑战赛第三赛道的技术方案。该任务要求基于视频内容识别道路场景中的64类原子活动。针对小目标、单体与群体对象区分困难以及模型过拟合等问题,我们提出多分支活动识别框架,分别处理不同物体类别及单体/群体识别任务,提升准确率。同时,设计多种模型集成策略,包括多帧采样序列、不同采样长度、多训练轮次及不同主干网络的融合。此外,提出一种原子活动识别的数据增强方法,通过翻转视频帧与道路拓扑结构,显著扩大样本空间,有效缓解过拟合。最终在ROAD++ Challenge 2024测试集上取得第一,达到69% mAP。

原文摘要 · Abstract (English)

This report presents our team's technical solution for participating in Track 3 of the 2024 ECCV ROAD++ Challenge. The task of Track 3 is atomic activity recognition, which aims to identify 64 types of atomic activities in road scenes based on video content. Our approach primarily addresses the challenges of small objects, discriminating between single object and a group of objects, as well as model overfitting in this task. Firstly, we construct a multi-branch activity recognition framework that not only separates different object categories but also the tasks of single object and object group recognition, thereby enhancing recognition accuracy. Subsequently, we develop various model ensembling strategies, including integrations of multiple frame sampling sequences, different frame sampling sequence lengths, multiple training epochs, and different backbone networks. Furthermore, we propose an atomic activity recognition data augmentation method, which greatly expands the sample space by flipping video frames and road topology, effectively mitigating model overfitting. Our methods rank first in the test set of Track 3 for the ROAD++ Challenge 2024, and achieve 69% mAP.

动作识别道路场景多分支数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。