用物理规律设计预训练任务,让传感器识活动更准且少标注。
PIM: Physics-Informed Multi-task Pre-training for Improving Inertial Sensor-Based Human Activity Recognition
- 基于人体运动的物理特性设计预训练任务,如速度、角度和对称性。
- 仅需每类2-8个标签样本,宏平均F1提升近10%,全量数据时提3%。
- 适合小样本或多传感器场景下的动作识别研究者使用。
基于深度学习的人体动作识别(HAR)依赖大量标注数据,获取成本高。自监督学习(SSL)通过掩码重建或多任务学习等预训练方法利用无标签数据,但现有方法多源自计算机视觉,忽视可穿戴传感器数据背后的物理机制。本文提出物理信息多任务预训练(PIM)框架,针对惯性测量单元(IMU)数据,基于人体运动的基本物理特性——如运动速度、运动角度及传感器布局对称性——构建预训练任务。对原始信号计算物理特征,作为自监督学习的前置任务。该方法使模型捕捉人体动作的本质物理属性,尤其适用于多传感器系统。在四个标准HAR数据集上的实验表明,本方法优于现有先进方法,包括数据增强和掩码重建,在仅需每类2至8个标注样本时,宏平均F1与准确率均提升近10%;在不减少训练数据量的情况下,性能提升达3%。
原文摘要 · Abstract (English)
Human activity recognition (HAR) with deep learning models relies on large amounts of labeled data, often challenging to obtain due to associated cost, time, and labor. Self-supervised learning (SSL) has emerged as an effective approach to leverage unlabeled data through pretext tasks, such as masked reconstruction and multitask learning with signal processing-based data augmentations, to pre-train encoder models. However, such methods are often derived from computer vision approaches that disregard physical mechanisms and constraints that govern wearable sensor data and the phenomena they reflect. In this paper, we propose a physics-informed multi-task pre-training (PIM) framework for IMU-based HAR. PIM generates pre-text tasks based on the understanding of basic physical aspects of human motion: including movement speed, angles of movement, and symmetry between sensor placements. Given a sensor signal, we calculate corresponding features using physics-based equations and use them as pretext tasks for SSL. This enables the model to capture fundamental physical characteristics of human activities, which is especially relevant for multi-sensor systems. Experimental evaluations on four HAR benchmark datasets demonstrate that the proposed method outperforms existing state-of-the-art methods, including data augmentation and masked reconstruction, in terms of accuracy and F1 score. We have observed gains of almost 10\% in macro f1 score and accuracy with only 2 to 8 labeled examples per class and up to 3% when there is no reduction in the amount of training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。