arXiv:2410.04980cs.CV2024-10被引 21

顶视图比斜视图更准,通用模型在婴儿上表现优于专用模型。

Comparison of marker-less 2D image-based methods for infant pose estimation

  • 用顶视图替代斜视图提升婴儿姿态估计精度。
  • 通用模型ViTPose在婴儿数据上表现最佳,重训练后效果显著提升。
  • 专用婴儿姿态模型泛化能力差,跨数据集使用需谨慎。

本研究比较了通用与婴儿专用姿态估计算法在基于视频的自动一般运动评估(GMA)中的性能,以及拍摄角度(传统斜视与顶视)的影响。基于75例4至26周婴儿自发运动的4500个标注视频帧进行评估,通过与人工标注的距离和关键点正确率(PCK)对比分析。结果表明:在婴儿数据上,经过成人数据训练的通用模型ViTPose表现最优;婴儿专用模型未带来性能提升。但将通用模型在本数据集上微调后,姿态估计准确率显著提高。顶视图的姿态估计精度显著优于斜视图,尤其在髋部关键点检测方面。此外,婴儿专用模型在不同婴儿数据集间泛化能力有限,提示应谨慎选择和使用未在其上训练的数据集上的专用模型。尽管传统GMA采用斜视,但顶视可大幅提升姿态估计效果,建议在自动化GMA研究中引入顶视拍摄方案。

原文摘要 · Abstract (English)

In this study we compare the performance of available generic- and infant-pose estimators for a video-based automated general movement assessment (GMA), and the choice of viewing angle for optimal recordings, i.e., conventional diagonal view used in GMA vs. top-down view. We used 4500 annotated video-frames from 75 recordings of infant spontaneous motor functions from 4 to 26 weeks. To determine which pose estimation method and camera angle yield the best pose estimation accuracy on infants in a GMA related setting, the distance to human annotations and the percentage of correct key-points (PCK) were computed and compared. The results show that the best performing generic model trained on adults, ViTPose, also performs best on infants. We see no improvement from using infant-pose estimators over the generic pose estimators on our infant dataset. However, when retraining a generic model on our data, there is a significant improvement in pose estimation accuracy. The pose estimation accuracy obtained from the top-down view is significantly better than that obtained from the diagonal view, especially for the detection of the hip key-points. The results also indicate limited generalization capabilities of infant-pose estimators to other infant datasets, which hints that one should be careful when choosing infant pose estimators and using them on infant datasets which they were not trained on. While the standard GMA method uses a diagonal view for assessment, pose estimation accuracy significantly improves using a top-down view. This suggests that a top-down view should be included in recording setups for automated GMA research.

姿态估计婴儿评估视频分析顶视图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。