用生物力学指导视频健身教练,让反馈更精准可懂。
From 3D Pose to Prose: Biomechanics-Grounded Vision--Language Coaching
- 分三阶段融合视觉与3D骨骼数据,聚焦关键关节动作
- 在QEVD-bio-fit-coach上提升文本准确率与评价得分
- 适合需要精确动作指导的健身应用开发者
我们提出BioCoach,一种基于生物力学的视觉-语言健身教练框架,从实时视频中提供个性化指导。该框架通过三阶段流程:针对特定动作的自由度选择器,聚焦显著关节;结构化生物力学上下文,结合个体体征、周期分析与约束条件;以及视觉-生物力学条件化的反馈模块,利用交叉注意力生成具体可操作的文本。采用参数高效训练,冻结视觉与语言主干网络,实现透明且个性化的推理。为支持学习与公平评估,我们在QEVD-fit-coach基础上引入生物力学导向反馈,构建QEVD-bio-fit-coach,并提出生物力学感知的LLM评判指标。BioCoach在QEVD-bio-fit-coach上取得显著提升,在词汇与判断指标上表现优异,同时保持时间触发准确性;在原始QEVD-fit-coach上,也提升了文本质量与正确性,且触发时间接近原模型,证明显式运动学与约束信息对实现精准、阶段感知的教练反馈至关重要。
原文摘要 · Abstract (English)
We present BioCoach, a biomechanics-grounded vision--language framework for fitness coaching from streaming video. BioCoach fuses visual appearance and 3D skeletal kinematics, through a novel three-stage pipeline: an exercise-specific degree-of-freedom selector that focuses analysis on salient joints; a structured biomechanical context that pairs individualized morphometrics with cycle and constraint analysis; and a vision--biomechanics conditioned feedback module that applies cross-attention to generate precise, actionable text. Using parameter-efficient training that freezes the vision and language backbones, BioCoach yields transparent, personalized reasoning rather than pattern matching. To enable learning and fair evaluation, we augment QEVD-fit-coach with biomechanics-oriented feedback to create QEVD-bio-fit-coach, and we introduce a biomechanics-aware LLM judge metric. BioCoach delivers clear gains on QEVD-bio-fit-coach across lexical and judgment metrics while maintaining temporal triggering; on the original QEVD-fit-coach, it improves text quality and correctness with near-parity timing, demonstrating that explicit kinematics and constraints are key to accurate, phase-aware coaching.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。