arXiv:2607.16350eess.SPcs.AI2026-07

无需标注数据,用新架构从传感器数据中自动学出强泛化能力的活动识别特征。

Joint-Embedding Predictive Architecture for Sensor-based Activity Recognition

论文配图:Joint-Embedding Predictive Architecture for Sensor-based Activity Recognition
图 1 · 摘自论文原文
  • 设计双层次时序编码器,捕捉局部窗口与长序列的动态特征。
  • 引入改进的VICReg正则化,提升预训练稳定性并防止表征坍塌。
  • 在少量数据的过渡动作上表现优于监督模型,适合小样本场景。

基于传感器的人体活动识别(HAR)在完全监督学习下取得了显著进展。然而,这些模型依赖大量标注数据,而数据收集与标注成本高昂。为此,本文提出一种面向传感器HAR的联合嵌入预测架构(HAR-JEPA),旨在从未标注数据中学习鲁棒且可泛化的表征。该框架采用编码器,显式建模单个时间窗口内的细粒度局部时序特征及相邻窗口间的长期时序序列。此外,我们提出改进的方差-不变性-协方差正则化(VICReg)目标函数,引入计算轻量的范数项以稳定预训练阶段,平衡方差、不变性与协方差约束,防止表征坍塌。在两个基准连续活动数据集上的实验表明,所提框架成功学习到高质量表征。尤其在少数类、高方差的过渡动作(如坐起、坐卧转换)上,其泛化性能显著优于传统监督学习,后者因支持样本不足易过拟合。

原文摘要 · Abstract (English)

Sensor-based human activity recognition (HAR) has achieved significant progressed in fully supervised learning settings. However, these supervised learning models rely on large amount of labeled data, which require labor-intensive collection and meticulous annotation. To address these challenges, this paper proposes a Joint Embedding Predictive Architecture framework tailored for sensor-based HAR, designed to learn robust and generalizable representations from unlabeled datasets. The proposed framework features an encoder designed to explicitly model both the fine-grained local temporal representations within individual window and the long-term temporal sequence of adjacent windows. Furthermore, we introduce an improved Variance-Invariance-Covariance Regularization (VICReg) objective function that incorporates computationally lightweight norm term to stabilize the JEPA pre-training phase. This term balances variance, invariance and covariance constraints to prevent representation collapse. The proposed HAR-JEPA framework is evaluated using two benchmark continuously performed activity datasets. The results show that high-quality representations are successfully learned by the proposed framework. Furthermore, the representations learned by HAR-JEPA demonstrates superior generalization on minority, high variance transitional activities such as sit-to-stand and sit-to-lie where supervised learning tend to overfit due to limited support.

活动识别自监督传感器表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。