arXiv:2605.14772cs.CVcs.GR2026-05

从视频中直接推断肌肉激活状态,实现视觉与生物力学的精准连接。

BioHuman: Learning Biomechanical Human Representations from Video

论文配图:BioHuman: Learning Biomechanical Human Representations from Video
图 1 · 摘自论文原文
  • 基于仿真构建视频-运动-肌激活同步数据集
  • 单目视频输入下联合预测动作与肌肉活动,准确率高
  • 适合运动分析、康复评估等需要内生生物力学的应用

理解人体运动的深层生物力学机制对运动分析、康复和损伤风险评估至关重要。然而,这一领域受限于缺乏大规模带生物力学标注的数据集,以及现有方法无法直接从视觉观察中推断内部生物力学状态。本文提出一种基于仿真的框架,从现有动作捕捉数据集中估算肌肉激活,构建了包含同步视频、动作和激活信息的大型数据集 BioHuman10M。在此基础上,我们设计了 BioHuman 模型,仅以单目视频为输入,即可联合预测人体运动与肌肉激活,有效连接视觉观测与内部生物力学状态。大量实验表明,该模型能准确重建运动学动作与肌肉活动,并在不同个体和动作上具有良好泛化能力。我们认为,本方法为基于视频的生物力学理解树立了新基准,为物理可解释的人体建模开辟了新路径。

原文摘要 · Abstract (English)

Understanding human motion beyond surface kinematics is crucial for motion analysis, rehabilitation, and injury risk assessment. However, progress in this domain is limited by the lack of large-scale datasets with biomechanical annotations, and by existing approaches that cannot directly infer internal biomechanical states from visual observations. In this paper, we introduce a simulation-based framework for estimating muscle activations from existing motion capture datasets, resulting in BioHuman10M, a large-scale dataset with synchronized video, motion, and activations. Building on BioHuman10M, we propose BioHuman, an end-to-end model that takes monocular video as input and jointly predicts human motion and muscle activations, effectively bridging visual observations and internal biomechanical states. Extensive experiments demonstrate that BioHuman enables accurate reconstruction of both kinematic motion and muscle activity, and generalizes across diverse subjects and motions. We believe our approach establishes a new benchmark for video-based biomechanical understanding and opens up new possibilities for physically grounded human modeling.

生物力学视频生成人体建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。