arXiv:2510.02264cs.CVcs.AI2025-10

用单目视频评估人体运动,对比惯性传感器表现。

Paving the Way Towards Kinematic Assessment Using Monocular Video: A Preclinical Benchmark of State-of-the-Art Deep-Learning-Based 3D Human Pose Estimators Against Inertial Sensors in Daily Living Activities

  • 用视频模型推算关节角度,与惯性传感器数据比对。
  • MotionAGFormer表现最佳,误差仅9.27°,相关性达0.86。
  • 适合远程康复、运动分析等场景的低成本方案参考。

机器学习与可穿戴传感器的进步为实验室外的人体运动捕捉提供了新可能。准确评估真实环境中的运动对远程医疗、运动科学和康复至关重要。本研究在包含13种临床相关日常活动的VIDIMU数据集上,将单目视频3D姿态估计模型(MotionAGFormer、MotionBERT、MMPose 2D-to-3D、NVIDIA BodyTrack)与五惯性测量单元(IMUs)数据进行对比,采用OpenSim逆运动学计算关节角,遵循Human3.6M数据集格式(17个关键点)。结果表明,MotionAGFormer表现最优,整体均方根误差为9.27°±4.80°,平均绝对误差为7.86°±4.18°,皮尔逊相关系数0.86±0.15,决定系数R²为0.67±0.28。研究确认视频与传感器方法均可用于非实验室场景,但各有成本、可及性与精度权衡。该工作为健康成人中视频模型的临床应用提供依据,并为远程监测系统设计提供指导。

原文摘要 · Abstract (English)

Advances in machine learning and wearable sensors offer new opportunities for capturing and analyzing human movement outside specialized laboratories. Accurate assessment of human movement under real-world conditions is essential for telemedicine, sports science, and rehabilitation. This preclinical benchmark compares monocular video-based 3D human pose estimation models with inertial measurement units (IMUs), leveraging the VIDIMU dataset containing a total of 13 clinically relevant daily activities which were captured using both commodity video cameras and five IMUs. During this initial study only healthy subjects were recorded, so results cannot be generalized to pathological cohorts. Joint angles derived from state-of-the-art deep learning frameworks (MotionAGFormer, MotionBERT, MMPose 2D-to-3D pose lifting, and NVIDIA BodyTrack) were evaluated against joint angles computed from IMU data using OpenSim inverse kinematics following the Human3.6M dataset format with 17 keypoints. Among them, MotionAGFormer demonstrated superior performance, achieving the lowest overall RMSE ($9.27°\pm 4.80°$) and MAE ($7.86°\pm 4.18°$), as well as the highest Pearson correlation ($0.86 \pm 0.15$) and the highest coefficient of determination $R^{2}$ ($0.67 \pm 0.28$). The results reveal that both technologies are viable for out-of-the-lab kinematic assessment. However, they also highlight key trade-offs between video- and sensor-based approaches including costs, accessibility, and precision. This study clarifies where off-the-shelf video models already provide clinically promising kinematics in healthy adults and where they lag behind IMU-based estimates while establishing valuable guidelines for researchers and clinicians seeking to develop robust, cost-effective, and user-friendly solutions for telehealth and remote patient monitoring.

动作评估单目视频惯性传感器远程医疗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。