从单目图像直接推断人体关节力矩,实现真实场景下的生物力学分析。
Learning human joint torques from pixels

- 基于视觉逆动力学框架,结合姿态预训练与时序推理恢复力矩。
- 在63,369帧数据上达到1.7612 N·m/kg的平均误差,提升超39%。
- 适合做运动分析、人机交互和无标记动作捕捉的研究者参考。
从视觉观测中估计人体关节力矩是将生物力学分析从受控实验室推向真实运动场景的关键一步。现有方法多依赖表面肌电图、运动捕捉标记、测力台或模拟模仿数据,难以直接应用于普通RGB图像。本文提出VID——一个基于视觉的逆动力学数据集与基准,用于直接从真实单目图像预测人体关节力矩。VID包含63,369个同步帧,涵盖真实人体图像、运动学标注、人体测量属性及OpenSim生成的动力学标签,提供真实图像与生物力学监督的配对数据。我们还定义了标准化评估协议,涵盖整体力矩估计、关节特异性分析与动作特异性预测。为建立强基线模型,提出VID-Network,融合姿态预训练的空间概率特征、标记回归与时序力矩推理,从图像序列中恢复关节力矩。在VID上的实验表明,VID-Network整体mPJE达1.7612 N·m/kg,优于最佳基线39.81%,并在所有关节类型和多数动作类别中取得最低误差。VID建立了首个面向视觉驱动的人体逆动力学实用基准,为非受限环境下生物力学推断研究奠定基础。
原文摘要 · Abstract (English)
Estimating human joint torques from visual observations is a key step toward bringing biomechanical analysis from controlled laboratories to real-world movement scenarios. Existing torque estimation methods typically depend on surface electromyography, motion-capture markers, force plates, or simulated imitation data, which limits their applicability to ordinary RGB images. In this work, we introduce VID, a vision-based inverse dynamics dataset and benchmark for predicting human joint torques directly from real monocular images. VID contains 63,369 synchronized frames with real human images, kinematic annotations, anthropometric attributes, and OpenSim-derived dynamic labels, providing paired visual and biomechanical supervision for real-image inverse dynamics. We further define a standardized evaluation protocol covering overall torque estimation, joint-specific analysis, and action-specific prediction. To establish a strong reference model, we propose VID-Network, which combines pose-pretrained spatial probabilistic features, marker regression, and temporal torque inference to recover joint torques from image sequences. Experiments on VID show that VID-Network achieves an overall mPJE of 1.7612 N$\cdot$m/kg, improving over the best compared baseline by 39.81\%, and obtains the lowest error across all evaluated joint types and most action categories. VID establishes a first practical benchmark for vision-driven human inverse dynamics and provides a foundation for studying biomechanical inference in less constrained environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。