轻量级无人机视角人体姿态估计模型,实时高效且精度显著提升。
FlyPose: Towards Robust Human Pose Estimation From Aerial Views
- 采用多数据集训练的轻量级顶向下架构,适应复杂空拍视角。
- 在UAV-Human数据集上2D姿态估计算法提升16.3 mAP,检测平均增益6.8 mAP。
- 可在Jetson Orin上实现20毫秒延迟,适合无人机实时部署。
无人飞行器(UAV)在包裹递送、交通监控、灾害响应和基础设施巡检等场景中日益靠近人类活动区域,安全可靠运行需从空中视角准确感知人体姿态与动作。然而低分辨率、陡峭视角及(自)遮挡对现有方法构成挑战,尤其在实时性要求高的应用中。本文提出FlyPose,一种专为航空影像设计的轻量级端到端人体姿态估计流程。通过多数据集联合训练,在Manipal-UAV、VisDrone、HIT-UAV及自建数据集上的人体检测平均提升6.8 mAP;在挑战性强的UAV-Human数据集上,2D姿态估计性能提升16.3 mAP。FlyPose在Jetson Orin AGX开发套件上推理延迟约20毫秒(含预处理),已成功部署于四旋翼无人机进行飞行实验。同时发布FlyPose-104——一个小型但具有挑战性的航空人体姿态数据集,包含从困难视角手动标注的样本:https://github.com/farooqhassaan/FlyPose。
原文摘要 · Abstract (English)
Unmanned Aerial Vehicles (UAVs) are increasingly deployed in close proximity to humans for applications such as parcel delivery, traffic monitoring, disaster response and infrastructure inspections. Ensuring safe and reliable operation in these human-populated environments demands accurate perception of human poses and actions from an aerial viewpoint. This perspective challenges existing methods with low resolution, steep viewing angles and (self-)occlusion, especially if the application demands realtime feasibile models. We train and deploy FlyPose, a lightweight top-down human pose estimation pipeline for aerial imagery. Through multi-dataset training, we achieve an average improvement of 6.8 mAP in person detection across the test-sets of Manipal-UAV, VisDrone, HIT-UAV as well as our custom dataset. For 2D human pose estimation we report an improvement of 16.3 mAP on the challenging UAV-Human dataset. FlyPose runs with an inference latency of ~20 milliseconds including preprocessing on a Jetson Orin AGX Developer Kit and is deployed onboard a quadrotor UAV during flight experiments. We also publish FlyPose-104, a small but challenging aerial human pose estimation dataset, that includes manual annotations from difficult aerial perspectives: https://github.com/farooqhassaan/FlyPose.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。