多视角相机提升骨骼数据精度,显著改善动作识别效果
Skarimva: Skeleton-based Action Recognition is a Multi-view Application
- 利用多视角三角测量生成更精确的3D骨骼数据
- 提升现有顶尖动作识别模型性能,验证数据质量是瓶颈
- 多相机方案性价比高,适合实际应用
人体动作识别在人机智能交互中至关重要。尽管当前研究主要聚焦于改进基于骨骼的动作识别机器学习算法,但输入骨骼数据的质量却未受到足够重视。本文表明,通过多摄像头视角三角测量生成更准确的3D骨骼数据,可显著提升当前先进动作识别模型的性能。这说明输入数据质量目前已成为模型表现的限制因素。基于此结果,认为在多数实际应用场景中,多相机系统的成本-收益比非常有利,因此未来基于骨骼的动作识别研究应将多视角应用作为标准配置。
原文摘要 · Abstract (English)
Human action recognition plays an important role when developing intelligent interactions between humans and machines. While there is a lot of active research on improving the machine learning algorithms for skeleton-based action recognition, not much attention has been given to the quality of the input skeleton data itself. This work demonstrates that by making use of multiple camera views to triangulate more accurate 3D~skeletons, the performance of state-of-the-art action recognition models can be improved significantly. This suggests that the quality of the input data is currently a limiting factor for the performance of these models. Based on these results, it is argued that the cost-benefit ratio of using multiple cameras is very favorable in most practical use-cases, therefore future research in skeleton-based action recognition should consider multi-view applications as the standard setup.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。