arXiv:2506.05321cs.LG2025-06被引 27

让传感器模型直接学不完整数据,无需补全也能精准建模。

LSM-2: Learning from Incomplete Wearable Sensor Data

  • 用可学习的掩码令牌同时处理真实缺失和人工缺失数据。
  • 在4000万小时数据上预训练,多任务表现最优。
  • 特别适合夜间生理信号等临床相关缺失场景。

基础模型虽推动了机器学习进步,但多依赖完整结构化数据。可穿戴传感器数据常存在大量缺失,给自监督学习(SSL)带来挑战。本文提出第二代大型传感器模型 LSM-2 与自适应继承掩码(AIM),一种新型 SSL 方法,可直接从不完整数据中学习鲁棒表征,无需显式插补。其核心创新在于使用可学习掩码令牌,同时建模真实存在的“继承”缺失与人为引入的缺失,使模型在推理时能稳健应对碎片化现实数据。在包含4000万小时全天候多模态传感器数据的大规模数据集上预训练后,采用 AIM 的 LSM-2 在分类、回归与生成建模等多种任务中均取得最佳性能。此外,该模型展现出优异的可扩展性,尤其在针对性缺失场景下仍保持高精度,反映出如夜间生物信号对高血压预测具有诊断价值等临床相关模式,使其成为真实可穿戴数据应用的更可靠选择。

原文摘要 · Abstract (English)

Foundation models, a cornerstone of recent advancements in machine learning, have predominantly thrived on complete and well-structured data. Wearable sensor data frequently suffers from significant missingness, posing a substantial challenge for self-supervised learning (SSL) models that typically assume complete data inputs. This paper introduces the second generation of Large Sensor Model (LSM-2) with Adaptive and Inherited Masking (AIM), a novel SSL approach that learns robust representations directly from incomplete data without requiring explicit imputation. AIM's core novelty lies in its use of learnable mask tokens to model both existing ("inherited") and artificially introduced missingness, enabling it to robustly handle fragmented real-world data during inference. Pre-trained on an extensive dataset of 40M hours of day-long multimodal sensor data, our LSM-2 with AIM achieves the best performance across a diverse range of tasks, including classification, regression and generative modeling. Furthermore, LSM-2 with AIM exhibits superior scaling performance, and critically, maintains high performance even under targeted missingness scenarios, reflecting clinically coherent patterns, such as the diagnostic value of nighttime biosignals for hypertension prediction. This makes AIM a more reliable choice for real-world wearable data applications.

传感器数据自监督学习缺失值处理可穿戴设备

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。