解决道路场景中时空行人检测的四大难题,夺冠表现优异。
First Place Solution to the ECCV 2024 ROAD++ Challenge @ ROAD++ Spatiotemporal Agent Detection 2024
- 设计双流检测模型,融合低光增强与特征融合提升弱光表现。
- 引入多分支框架和预训练微调策略,改善类别不平衡与细粒度分类。
- 在低光、极端大小目标等复杂场景下实现30.82%视频mAP,排名第一。
本文介绍我们团队在2024年ECCV ROAD++挑战赛第1赛道中的解决方案。该任务为时空行人检测,旨在连续视频帧中构建“行人轨迹管”。针对极端尺寸目标、低光照、类别不平衡及细粒度分类等挑战,我们提出:引入极端尺寸目标检测头以提升大/小目标检测性能;设计双流检测模型,包含低光增强分支与特征融合模块,增强弱光场景下的检测能力;构建多分支检测框架,结合预训练与微调策略优化类别不平衡与细粒度分类问题;同时采用数据增强、损失函数改进与上采样优化。最终在测试集上取得30.82%的平均视频mAP,位列第一。
原文摘要 · Abstract (English)
This report presents our team's solutions for the Track 1 of the 2024 ECCV ROAD++ Challenge. The task of Track 1 is spatiotemporal agent detection, which aims to construct an "agent tube" for road agents in consecutive video frames. Our solutions focus on the challenges in this task, including extreme-size objects, low-light scenarios, class imbalance, and fine-grained classification. Firstly, the extreme-size object detection heads are introduced to improve the detection performance of large and small objects. Secondly, we design a dual-stream detection model with a low-light enhancement stream to improve the performance of spatiotemporal agent detection in low-light scenes, and the feature fusion module to integrate features from different branches. Subsequently, we develop a multi-branch detection framework to mitigate the issues of class imbalance and fine-grained classification, and we design a pre-training and fine-tuning approach to optimize the above multi-branch framework. Besides, we employ some common data augmentation techniques, and improve the loss function and upsampling operation. We rank first in the test set of Track 1 for the ROAD++ Challenge 2024, and achieve 30.82% average video-mAP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。