arXiv:2410.16037cs.CV2024-10

提升交通场景下多标签动作识别的准确率,关键在鲁棒特征与注意力机制优化。

Improving the Multi-label Atomic Activity Recognition by Robust Visual Feature and Advanced Attention @ ROAD++ Atomic Activity Recognition 2024

  • 采用固定采样策略与多骨干网络提取鲁棒视觉特征。
  • 引入动作槽模型,测试集上mAP达58%,比基线高4%。
  • 融合多模型结果加权输出,适合真实交通场景应用。

ROAD++ Track3 提出了一个交通场景下的多标签原子动作识别任务,可标准化为64类多标签视频动作识别。该任务中,视觉特征提取的鲁棒性仍是关键挑战,直接影响模型性能与泛化能力。为此,本团队从数据处理、模型训练和后处理三方面进行优化:首先,选择合适的分辨率与视频采样策略,并在验证集和测试集上采用固定采样;其次,在模型训练中,选用多种视觉骨干网络提取特征,并引入动作槽模型,在训练集和验证集上训练,于测试集推理;最后,后处理阶段通过融合不同模型的优劣进行加权融合,最终在测试集上获得58%的mAP,较挑战基线提升4%。

原文摘要 · Abstract (English)

Road++ Track3 proposes a multi-label atomic activity recognition task in traffic scenarios, which can be standardized as a 64-class multi-label video action recognition task. In the multi-label atomic activity recognition task, the robustness of visual feature extraction remains a key challenge, which directly affects the model performance and generalization ability. To cope with these issues, our team optimized three aspects: data processing, model and post-processing. Firstly, the appropriate resolution and video sampling strategy are selected, and a fixed sampling strategy is set on the validation and test sets. Secondly, in terms of model training, the team selects a variety of visual backbone networks for feature extraction, and then introduces the action-slot model, which is trained on the training and validation sets, and reasoned on the test set. Finally, for post-processing, the team combined the strengths and weaknesses of different models for weighted fusion, and the final mAP on the test set was 58%, which is 4% higher than the challenge baseline.

动作识别多标签视觉特征交通场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。