arXiv:2410.19818eess.SPcs.AI2024-10NeurIPS被引 44

首个统一预训练的运动时序模型,跨设备、跨动作泛化能力极强。

UniMTS: Unified Pre-training for Motion Time Series

  • 用对比学习对齐时序数据与大模型生成的文本描述,学习语义表示。
  • 在18个基准数据集上零样本性能超基线340%,少样本和全样本也显著领先。
  • 适用于隐私敏感场景,无需大规模真实数据,适合医疗与可穿戴设备应用。

来自手机和智能手表等可穿戴设备的运动时序数据具有低功耗、持续运行的特点,在医疗、自动化、物联网及AR/XR领域应用广泛。但受安全与隐私限制,构建大规模运动时序数据集困难,阻碍了人类活动分析预训练模型的发展。现有模型通常在相同数据集上训练测试,导致在设备位置、佩戴方向和活动类型变化下泛化能力差。本文提出UniMTS,首个统一预训练的运动时序模型,能跨设备隐含因素与活动类型实现良好泛化。我们采用对比学习框架,将运动时序数据与大语言模型增强的文本描述对齐,使模型学习时序数据语义。由于缺乏大规模真实数据,我们从已有运动骨架数据中合成覆盖全关节的时序数据。利用时空图网络捕捉关节间关系,提升不同设备位置下的泛化能力。设计旋转不变性增强策略,使模型对设备佩戴方向变化不敏感。实验显示,该模型在18个运动时序分类基准数据集上表现优异,零样本设置下超越最佳基线340%,少样本设置下提升16.3%,全样本设置下提升9.2%。

原文摘要 · Abstract (English)

Motion time series collected from mobile and wearable devices such as smartphones and smartwatches offer significant insights into human behavioral patterns, with wide applications in healthcare, automation, IoT, and AR/XR due to their low-power, always-on nature. However, given security and privacy concerns, building large-scale motion time series datasets remains difficult, preventing the development of pre-trained models for human activity analysis. Typically, existing models are trained and tested on the same dataset, leading to poor generalizability across variations in device location, device mounting orientation and human activity type. In this paper, we introduce UniMTS, the first unified pre-training procedure for motion time series that generalizes across diverse device latent factors and activities. Specifically, we employ a contrastive learning framework that aligns motion time series with text descriptions enriched by large language models. This helps the model learn the semantics of time series to generalize across activities. Given the absence of large-scale motion time series data, we derive and synthesize time series from existing motion skeleton data with all-joint coverage. Spatio-temporal graph networks are utilized to capture the relationships across joints for generalization across different device locations. We further design rotation-invariant augmentation to make the model agnostic to changes in device mounting orientations. Our model shows exceptional generalizability across 18 motion time series classification benchmark datasets, outperforming the best baselines by 340% in the zero-shot setting, 16.3% in the few-shot setting, and 9.2% in the full-shot setting.

运动时序预训练对比学习可穿戴

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。