arXiv:2506.09345cs.CV2025-06被引 4

通过数据增强与多模态融合,实现高精度动作识别

An Effective End-to-End Solution for Multimodal Action Recognition

  • 优化数据增强并融合多源RGB数据预训练模型
  • 用2D CNN+TSM实现高效时空特征提取,媲美3D CNN
  • 集成SWA、Ensemble和TTA提升预测性能,适合竞赛应用

近年来,多模态任务凭借丰富的信息显著推动了动作识别的发展。然而,由于三模态数据稀缺,三模态动作识别面临诸多挑战。为此,我们提出了一套全面的多模态动作识别解决方案,有效利用多模态信息。首先,通过优化数据增强技术扩展现有数据,扩大训练规模;同时引入更多RGB数据预训练主干网络,借助迁移学习使其更适应新任务。其次,利用2D CNN提取多模态空间特征,并结合时间移位模块(TSM)实现媲美3D CNN的多模态时空特征提取,同时提升计算效率。此外,采用随机权重平均(SWA)、集成学习和测试时增强(TTA)等通用预测增强方法,融合同一架构不同训练阶段及不同架构模型的知识,从多角度预测动作,充分挖掘目标信息。最终,在竞赛排行榜上取得Top-1准确率99%、Top-5准确率100%的成绩,验证了该方案的优越性。

原文摘要 · Abstract (English)

Recently, multimodal tasks have strongly advanced the field of action recognition with their rich multimodal information. However, due to the scarcity of tri-modal data, research on tri-modal action recognition tasks faces many challenges. To this end, we have proposed a comprehensive multimodal action recognition solution that effectively utilizes multimodal information. First, the existing data are transformed and expanded by optimizing data enhancement techniques to enlarge the training scale. At the same time, more RGB datasets are used to pre-train the backbone network, which is better adapted to the new task by means of transfer learning. Secondly, multimodal spatial features are extracted with the help of 2D CNNs and combined with the Temporal Shift Module (TSM) to achieve multimodal spatial-temporal feature extraction comparable to 3D CNNs and improve the computational efficiency. In addition, common prediction enhancement methods, such as Stochastic Weight Averaging (SWA), Ensemble and Test-Time augmentation (TTA), are used to integrate the knowledge of models from different training periods of the same architecture and different architectures, so as to predict the actions from different perspectives and fully exploit the target information. Ultimately, we achieved the Top-1 accuracy of 99% and the Top-5 accuracy of 100% on the competition leaderboard, demonstrating the superiority of our solution.

动作识别多模态模型融合竞赛优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。