arXiv:2608.20969cs.CV2026-08

用结构化运动知识图谱提升多模态步态分析的可解释性与泛化能力

Kinematic Knowledge Maps for Pattern Alignment: Structured Latent Representational Learning in Multimodal Gait Analysis

  • 构建固定索引的运动知识图谱,融合绝对运动、骨架配置与关节关联特征
  • 在1858例数据上实现0.972的外部AUC,优于单模态和后期拼接方法
  • 可直接映射到特定运动阶段与骨骼指标,提供可验证的可解释性

多模态临床AI受限于输入弱对齐及缺乏领域可解释表示,尤其在处理密集视频流、结构化时间序列和基于模板的运动学文本时。本文提出ScoliDetect,一种从单目步态视频进行青少年特发性脊柱侧弯筛查的可解释框架,核心为运动知识图谱(KKM)与基于序列姿态静态的互补模板化运动学文本。KKM是一种固定索引的结构化表示,编码绝对运动、自骨架配置及关节-关节信号相关性特征,支持锚点参考的多模态融合与因子级解释。通过双向交叉注意力与隐空间瓶颈聚合整合视频、KKM与模板化运动学文本。在多中心队列(剔除后n=1,858)中,预设监督消融实验显示,基于KKM的多模态融合优于单模态模型与晚期拼接。经分阶段训练流程,架构选定后进行三模态对比预训练作为表征初始化,使外部ROC-AUC由0.961提升至0.972。此外,KKM的结构特性提供了内在的因子级归因,可直接映射至具体运动阶段与骨骼索引,实现可验证的可解释性。结果表明,将显式拓扑结构嵌入潜在空间能显著增强多模态模式分析系统的泛化性与可解释性。

原文摘要 · Abstract (English)

Multimodal clinical AI is limited by weakly aligned inputs and the absence of domain-specific interpretable representations, particularly when learning from dense video stream, structured time-series, and template-based kinematic text. Here we present ScoliDetect, an explainable framework for adolescent idiopathic scoliosis screening from monocular gait video, built around a kinematic knowledge map (KKM) and complementary template-based kinematic text derived from per-sequence pose statics. KKM is a fixed-index structured representation that encodes gait features across absolute motion, self-skeleton configuration and joint-joint signal correlation, providing anchor-referenced multimodal fusion and factor-level interpretation. We integrate video, KKM, and template-based kinematic text through bidirectional cross-attention with latent-bottleneck aggregation. In a multicenter cohort (n = 1,858 after exclusions), prespecified supervised ablations on an external screening cohort show that KKM-mediated multimodal fusion outperforms unimodal models and late concatenation. Under a staged training protocol, trimodal contrastive pretraining is applied after architecture selection as representation initialization, improving external ROC-AUC from 0.961 to 0.972. Furthermore, the structured nature of the KKM provides inherent, factor-level attributions mapped directly to specific kinematic phases and skeletal indices, offering verifiable interpretability. The results demonstrate that embedding explicit structural topologies into latent spaces significantly enhances both the generalization and explainability of multimodal pattern analysis systems.

多模态学习可解释性运动分析医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。