arXiv:2603.13317cs.LGcs.HC2026-03被引 1

用文字描述步态数据,大模型分类效果不如传统方法。

Evaluating Large Language Models for Gait Classification Using Text-Encoded Kinematic Waveforms

  • 将步态数据转为文本序列输入大模型,尝试实现可解释分类。
  • 最佳大模型(GPT-5)多类MCC达0.70,低于传统KNN的0.88。
  • 高置信度预测时表现提升,适合辅助探索而非临床诊断。

背景:机器学习虽能提升步态分析效率,但临床应用常受限于可解释性不足。大语言模型(LLMs)在结构化运动数据上或可提供解释与置信度输出。本研究评估通用大模型在将连续步态运动学数据以文本数字序列表示后,能否实现步态分类,并对比其与传统机器学习方法的表现。方法:20名参与者完成7种步态模式,采集下肢运动学数据。采用监督KNN分类器与无类别依赖的One-Class SVM(OCSVM),对比零样本大模型(GPT-5、GPT-5-mini、GPT-4.1、o4-mini)。使用留一被试者交叉验证(LOSO)。大模型在有无显式参考步态统计信息条件下均测试。结果:监督KNN表现最佳(多类MCC=0.88)。最优大模型(GPT-5)在引入参考信息后,多类MCC为0.70,二分类MCC为0.68,优于无类别独立的OCSVM(二分类MCC=0.60)。大模型性能高度依赖显式参考信息与自评置信度;当仅保留高置信预测时,多类MCC提升至0.83。值得注意的是,计算高效的o4-mini模型表现与更大模型相当。结论:将连续运动学波形编码为文本数值标记后,通用大模型即使有参考引导,仍无法达到监督分类器在精确步态分类上的水平,更适合作为需谨慎人工审校的探索性工具,而非诊断系统。

原文摘要 · Abstract (English)

Background: Machine learning (ML) enhances gait analysis but often lacks the level of interpretability desired for clinical adoption. Large Language Models (LLMs) may offer explanatory capabilities and confidence-aware outputs when applied to structured kinematic data. This study therefore evaluated whether general-purpose LLMs can classify continuous gait kinematics when represented as textual numeric sequences and how their performance compares to conventional ML approaches. Methods: Lower-body kinematics were recorded from 20 participants performing seven gait patterns. A supervised KNN classifier and a class-independent One-Class SVM (OCSVM) were compared against zero-shot LLMs (GPT-5, GPT-5-mini, GPT-4.1, and o4-mini). Models were evaluated using Leave-One-Subject-Out (LOSO) cross-validation. LLMs were tested both with and without explicit reference gait statistics. Results: The supervised KNN achieved the highest performance (multiclass Matthews Correlation Coefficient, MCC = 0.88). The best-performing LLM (GPT-5) with reference grounding achieved a multiclass MCC of 0.70 and a binary MCC of 0.68, outperforming the class-independent OCSVM (binary MCC = 0.60). Performance of the LLM was highly dependent on explicit reference information and self-rated confidence; when restricted to high-confidence predictions, multiclass MCC increased to 0.83 on the filtered subset. Notably, the computationally efficient o4-mini model performed comparably to larger models. Conclusion: When continuous kinematic waveforms were encoded as textual numeric tokens, general-purpose LLMs, even with reference grounding, did not match supervised multiclass classifiers for precise gait classification and are better regarded as exploratory systems requiring cautious, human-guided interpretation rather than diagnostic use.

大模型步态识别可解释性运动分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。