arXiv:2604.15665cs.CVcs.PF2026-04

让单目3D步态分析在普通电脑上高效运行,无需显卡。

CPU Optimization of a Monocular 3D Biomechanics Pipeline for Low-Resource Deployment

  • 重构模型初始化,消除硬盘读写瓶颈,提升多核并行效率。
  • 处理速度提升2.47倍,总耗时减少59.6%,启动延迟降低4.6倍。
  • 结果与原版高度一致,适合临床和运动场景的低成本部署。

从单目视频中实现无标记3D运动分析,可为临床和运动场景提供可及的生物力学评估。然而,多数研究级流程依赖GPU加速,限制了在消费级硬件和低资源环境中的部署。本文针对基于MonocularBiomechanics框架的单目3D生物力学流程,优化其在纯CPU环境下的执行效率。通过面向性能分析的系统优化,包括模型初始化重构、消除磁盘I/O串行化以及改进CPU并行策略。在搭载AMD Ryzen 7 9700X CPU的消费级工作站上实验表明,处理吞吐量提升2.47倍,总运行时间减少59.6%,初始化延迟降低4.6倍。尽管进行了上述调整,生物力学输出仍与基准版本高度一致(平均关节角偏差0.35°,相关系数r=0.998)。结果表明,研究级视觉生物力学流程可在通用CPU硬件上部署,实现可扩展的运动评估。

原文摘要 · Abstract (English)

Markerless 3D movement analysis from monocular video enables accessible biomechanical assessment in clinical and sports settings. However, most research-grade pipelines rely on GPU acceleration, limiting deployment on consumer-grade hardware and in low-resource environments. In this work, we optimize a monocular 3D biomechanics pipeline derived from the MonocularBiomechanics framework for efficient CPU-only execution. Through profiling-driven system optimization, including model initialization restructuring, elimination of disk I/O serialization, and improved CPU parallelization. Experiments on a consumer workstation (AMD Ryzen 7 9700X CPU) show a 2.47x increase in processing throughput and a 59.6\% reduction in total runtime, with initialization latency reduced by 4.6x. Despite these changes, biomechanical outputs remain highly consistent with the baseline implementation (mean joint-angle deviation 0.35$^\circ$, $r=0.998$). These results demonstrate that research-grade vision-based biomechanics pipelines can be deployed on commodity CPU hardware for scalable movement assessment.

3D步态单目视觉边缘计算生物力学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。