arXiv:2504.11726cs.LGcs.AI2025-04被引 1

用海量无标签传感器数据提升用户感知精度,少量标注即可达高准确率。

Saga: Capturing Multi-granularity Semantics from Massive Unlabelled IMU Data for User Perception

  • 利用大规模无标签IMU数据预训练多粒度语义特征模型
  • 仅需每类100样本即达90%以上准确率,媲美万级标注数据模型
  • 适合作为移动端用户感知任务的轻量高效解决方案

惯性测量单元(IMU)广泛应用于活动识别与用户认证等移动感知任务,通常需要大量标注数据训练模型。然而,由于原始IMU数据难以理解且缺乏真实标签,对海量IMU数据中的微小动作进行标注极为困难。本文提出一种细粒度用户感知方法Saga,仅需少量标注数据即可实现卓越的感知准确率。其核心思想是首先利用海量无标签IMU数据中蕴含的多粒度语义信息,预训练一个主干特征提取模型;同时,针对具体下游任务,采用贝叶斯优化确定不同语义层级预训练任务的最优权重。我们在五款典型手机上实现Saga,并在三个典型任务和三个IMU数据集上进行评估。结果表明,当每类仅使用约100个训练样本时,Saga可达到超过90%的准确率,媲美在超万样本上训练的完整模型,且无额外系统开销。

原文摘要 · Abstract (English)

Inertial measurement units (IMUs), have been prevalently used in a wide range of mobile perception applications such as activity recognition and user authentication, where a large amount of labelled data are normally required to train a satisfactory model. However, it is difficult to label micro-activities in massive IMU data due to the hardness of understanding raw IMU data and the lack of ground truth. In this paper, we propose a novel fine-grained user perception approach, called Saga, which only needs a small amount of labelled IMU data to achieve stunning user perception accuracy. The core idea of Saga is to first pre-train a backbone feature extraction model, utilizing the rich semantic information of different levels embedded in the massive unlabelled IMU data. Meanwhile, for a specific downstream user perception application, Bayesian Optimization is employed to determine the optimal weights for pre-training tasks involving different semantic levels. We implement Saga on five typical mobile phones and evaluate Saga on three typical tasks on three IMU datasets. Results show that when only using about 100 training samples per class, Saga can achieve over 90% accuracy of the full-fledged model trained on over ten thousands training samples with no additional system overhead.

用户感知IMU数据少样本学习预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。