让模型同时学动作和手势,提升识别效果与泛化能力
Multi-task Learning For Joint Action and Gesture Recognition
- 用多任务学习统一建模动作与手势识别
- 在多个数据集上双任务表现优于单任务模型
- 适合需要联合理解身体与手部动作的场景
在实际应用中,计算机视觉任务常需同时处理。多任务学习通常通过联合训练单一深度神经网络来学习共享表征,提升效率并增强泛化能力。尽管动作与手势识别密切相关,且分别关注身体与手部运动,现有最先进方法仍将其分开处理。本文表明,采用多任务学习框架进行动作与手势识别,可利用两者间的协同效应,获得更高效、更鲁棒、更具泛化性的视觉表征。在多个动作与手势数据集上的大量实验表明,将两者统一于单一架构中,相比各自的单任务学习模型,能实现更好的性能表现。
原文摘要 · Abstract (English)
In practical applications, computer vision tasks often need to be addressed simultaneously. Multitask learning typically achieves this by jointly training a single deep neural network to learn shared representations, providing efficiency and improving generalization. Although action and gesture recognition are closely related tasks, since they focus on body and hand movements, current state-of-the-art methods handle them separately. In this paper, we show that employing a multi-task learning paradigm for action and gesture recognition results in more efficient, robust and generalizable visual representations, by leveraging the synergies between these tasks. Extensive experiments on multiple action and gesture datasets demonstrate that handling actions and gestures in a single architecture can achieve better performance for both tasks in comparison to their single-task learning variants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。