arXiv:2411.03129physics.bio-phcs.CV2024-11

用自监督与运动增强提升步态识别疾病精度,仅需1%标注数据

MA^2: A Self-Supervised and Motion Augmenting Autoencoder for Gait-Based Automatic Disease Detection

  • 自监督训练+运动数据增强,减少人工标注依赖
  • 在1%标注数据下达90.91%准确率,帕金森数据集泛化达78.57%
  • 适合医疗诊断、小样本疾病识别研究者使用

足底反作用力(GRF)是地面作用于接触物体的力。基于GRF的自动疾病检测(ADD)是一种新兴的医学诊断方法,旨在通过深度学习识别不同步态压力对应的疾病模式。尽管现有方法可节省医生诊断时间,但深度模型训练仍面临大量受试者步态数据标注成本高的问题。此外,统一基准GRF数据集上的模型准确率与在可扩展步态数据集上的泛化能力仍有待提升。为此,本文提出MA2,一种基于GRF的自监督运动增强自编码器,将ADD任务建模为编码器-解码器范式。编码器引入三层1D卷积提取特征,并设计掩码生成器随机遮蔽特征序列,以最大化模型捕捉高层次判别性内在表示的能力;解码器则利用该信息重建原始输入序列并计算重建损失以优化网络。自编码器主干采用多头自注意力机制,能全局捕捉输入特征的上下文信息,而非局限于局部邻域。大量实验表明,MA2在仅1%标注病理GRF样本下达到90.91%的准确率,在可扩展帕金森病数据集上展现78.57%的泛化性能,表现优于现有方法。

原文摘要 · Abstract (English)

Ground reaction force (GRF) is the force exerted by the ground on a body in contact with it. GRF-based automatic disease detection (ADD) has become an emerging medical diagnosis method, which aims to learn and identify disease patterns corresponding to different gait pressures based on deep learning methods. Although existing ADD methods can save doctors time in making diagnoses, training deep models still struggles with the cost caused by the labeling engineering for a large number of gait diagnostic data for subjects. On the other hand, the accuracy of the deep model under the unified benchmark GRF dataset and the generalization ability on scalable gait datasets need to be further improved. To address these issues, we propose MA2, a GRF-based self-supervised and motion augmenting auto-encoder, which models the ADD task as an encoder-decoder paradigm. In the encoder, we introduce an embedding block including the 3-layer 1D convolution for extracting the token and a mask generator to randomly mask out the sequence of tokens to maximize the model's potential to capture high-level, discriminative, intrinsic representations. whereafter, the decoder utilizes this information to reconstruct the pixel sequence of the origin input and calculate the reconstruction loss to optimize the network. Moreover, the backbone of an auto-encoder is multi-head self-attention that can consider the global information of the token from the input, not just the local neighborhood. This allows the model to capture generalized contextual information. Extensive experiments demonstrate MA2 has SOTA performance of 90.91% accuracy on 1% limited pathological GRF samples with labels, and good generalization ability of 78.57% accuracy on scalable Parkinson disease dataset.

步态识别自监督学习疾病检测小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。