arXiv:2603.08564cs.CV2026-03被引 2

用视觉+语言+生物力学三模态提升步态分析的可解释性

BioGait-VLM: A Tri-Modal Vision-Language-Biomechanics Framework for Interpretable Clinical Gait Assessment

  • 融合视频、语言和生物力学信息,避免模型依赖环境线索
  • 在8类步态疾病上达到顶尖识别准确率,且无数据泄露
  • 专家验证显示结果更符合临床逻辑,适合医疗场景应用

基于视频的临床步态分析常因模型过拟合环境特征而泛化能力差,未能捕捉病理性运动。为此,我们提出BioGait-VLM,一种三模态视觉-语言-生物力学框架,用于可解释的临床步态评估。不同于标准视频编码器,该架构引入时序证据提炼分支以捕捉节律动态,并设计生物力学分词分支,将3D骨骼序列映射为与语言对齐的语义标记,使模型能独立于视觉捷径显式推理关节力学。为确保严谨评估,我们在公开GAVD数据集基础上,加入高保真退行性颈椎髓病(DCM)队列,构建统一的8类分类体系,并采用严格的人体无关划分协议防止数据泄露。在此设置下,BioGait-VLM实现领先识别准确率。此外,盲法专家研究证实,生物力学标记显著提升临床合理性与证据依据性,为透明、隐私保护的步态评估提供新路径。

原文摘要 · Abstract (English)

Video-based Clinical Gait Analysis often suffers from poor generalization as models overfit environmental biases instead of capturing pathological motion. To address this, we propose BioGait-VLM, a tri-modal Vision-Language-Biomechanics framework for interpretable clinical gait assessment. Unlike standard video encoders, our architecture incorporates a Temporal Evidence Distillation branch to capture rhythmic dynamics and a Biomechanical Tokenization branch that projects 3D skeleton sequences into language-aligned semantic tokens. This enables the model to explicitly reason about joint mechanics independent of visual shortcuts. To ensure rigorous benchmarking, we augment the public GAVD dataset with a high-fidelity Degenerative Cervical Myelopathy (DCM) cohort to form a unified 8-class taxonomy, establishing a strict subject-disjoint protocol to prevent data leakage. Under this setting, BioGait-VLM achieves state-of-the-art recognition accuracy. Furthermore, a blinded expert study confirms that biomechanical tokens significantly improve clinical plausibility and evidence grounding, offering a path toward transparent, privacy-enhanced gait assessment.

步态分析三模态可解释性生物力学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。