仅用单摄像头视频预测关节受力,无需设备或建模。
From Pixels to Newtons: Predicting In Vivo Joint Contact Forces from Monocular Video

- 从单视角视频中恢复人体姿态,用Transformer模型直接输出三维关节力。
- 髋膝关节预测误差分别达0.32±0.08 BW和0.23±0.03 BW,媲美专业仿真。
- 无需标注动作类别,可直接在原始视频上端到端推断,适合临床应用。
关节接触力影响假体寿命、软骨健康及康复效果,但目前仅能通过植入式设备在极少数患者中测量。本文提出一种无需物理模型的管道,仅凭未标定的单目视频即可预测髋膝关节的瞬时三维接触力:无需标记点、测力台、肌电图、个体影像或肌肉骨骼模型。每帧恢复参数化身体网格,提取运动学特征,由变压器模型解码为力值,其姿态流在每一层自适应地融合身体形态、关节类型、侧别、活动文本及自监督视频标记(V-JEPA 2),统一建模髋膝。在26名患者、25类活动的OrthoLoad数据库中,采用留一患者交叉验证,该方法在髋关节(0.32±0.08 BW RMSE)与膝关节(0.23±0.03 BW)上的精度媲美个体化肌肉骨骼模拟,并可分辨小于步态再训练和骨关节炎进展报告的力变化。在独立植入式队列上零样本应用,表现优于或匹敌已有方法。即使无标注活动标签,仅靠视频特征仍保持精度,支持对原始视频端到端推理。基于预测器生成的运动先验可产生生物力学合理的变体,降低峰值载荷,复现预测模拟文献中的策略。该管道确立了未标定单目视频作为关节负荷估计的有效模态,为回顾性分析临床录像、初级筛查及居家康复监测开辟路径。
原文摘要 · Abstract (English)
Joint contact forces govern implant longevity, cartilage health, and rehabilitation outcomes, shaping who develops osteoarthritis, who recovers well from joint replacement, and who benefits from biomechanical interventions. Yet they remain measurable only invasively, in a few dozen patients with instrumented implants. I present a physics-free pipeline to predict instantaneous 3D hip and knee contact forces from an uncalibrated monocular video: no markers, force plates, electromyography, subject-specific imaging, or musculoskeletal model. Parametric body meshes are recovered per frame, encoded as kinematic features, and decoded into forces by a transformer whose pose stream is adaptively modulated at every layer by body shape, joint, side, activity text, and self-supervised video tokens (V-JEPA 2), unifying hip and knee in a single model. Under leave-one-subject-out cross-validation across 26 patients and 25 activity categories from the in vivo OrthoLoad database, the pipeline matches the accuracy of subject-specific musculoskeletal simulations ($0.32 \pm 0.08$ BW RMSE for hip; $0.23 \pm 0.03$ BW for knee) and resolves peak force changes smaller than those reported for gait retraining and osteoarthritis progression. Applied zero-shot to an independent instrumented cohort, it rivals or outperforms prior published methods. Even without curated activity labels, video features alone preserve accuracy and enable end-to-end inference on raw footage. Driven by the predictor, a generative motion prior produces biomechanically plausible variants with reduced peak loading, rediscovering strategies from the predictive simulation literature. This pipeline establishes uncalibrated monocular video as a viable modality for estimating joint loading, opening a path toward retrospective analysis of archived clinical recordings, primary-care screening, and at-home rehabilitation tracking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。