用惯性传感器与视频联合预训练,提升异常动作识别的泛化能力。
Improving Out-of-distribution Human Activity Recognition via IMU-Video Cross-modal Representation Learning
- 通过无标签的IMU-视频数据进行跨模态自监督预训练
- 在帕金森患者等外部数据集上零样本/少样本性能超越现有方法
- 适合远程医疗中动态环境下的动作识别任务
基于可穿戴惯性传感器的人体动作识别(HAR)在远程健康监测中至关重要。对于运动障碍患者,居家环境中检测异常动作可实现治疗持续优化并及时提醒照护者。现有机器学习方法多依赖特定应用场景标签,难以泛化到不同环境或人群的数据。为此,我们提出一种新的跨模态自监督预训练方法,利用大规模未标注的IMU-视频数据学习通用表征,并在外部分布(OOD)IMU数据集上验证其优越泛化能力,包括帕金森病患者数据集。结果表明,该方法在零样本和少样本评估下均优于当前最优的IMU-视频预训练及仅使用IMU的预训练方案。研究表明,在高动态信号如IMU数据中,跨模态预训练是获取通用表征的有效工具。相关代码已开源:https://github.com/scheshmi/IMU-Video-OOD-HAR。
原文摘要 · Abstract (English)
Human Activity Recognition (HAR) based on wearable inertial sensors plays a critical role in remote health monitoring. In patients with movement disorders, the ability to detect abnormal patient movements in their home environments can enable continuous optimization of treatments and help alert caretakers as needed. Machine learning approaches have been proposed for HAR tasks using Inertial Measurement Unit (IMU) data; however, most rely on application-specific labels and lack generalizability to data collected in different environments or populations. To address this limitation, we propose a new cross-modal self-supervised pretraining approach to learn representations from large-sale unlabeled IMU-video data and demonstrate improved generalizability in HAR tasks on out of distribution (OOD) IMU datasets, including a dataset collected from patients with Parkinson's disease. Specifically, our results indicate that the proposed cross-modal pretraining approach outperforms the current state-of-the-art IMU-video pretraining approach and IMU-only pretraining under zero-shot and few-shot evaluations. Broadly, our study provides evidence that in highly dynamic data modalities, such as IMU signals, cross-modal pretraining may be a useful tool to learn generalizable data representations. Our software is available at https://github.com/scheshmi/IMU-Video-OOD-HAR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。