arXiv:2506.09108cs.LGcs.AI2025-06NeurIPS被引 64

用自然语言理解可穿戴传感器数据,构建了超大规模数据集。

SensorLM: Learning the Language of Wearable Sensors

  • 设计分层生成管道,从传感器数据提取统计、结构和语义信息。
  • 建成超5970万小时数据集,覆盖10.3万人,是目前最大规模。
  • 支持零样本识别与跨模态检索,适合健康分析与动作识别场景。

我们提出SensorLM,一类可理解可穿戴传感器数据的传感-语言基础模型。尽管可穿戴设备应用广泛,但将传感器数据与自然语言对齐仍面临挑战,主要因缺乏真实世界中未加标注的数据里成对的丰富传感器-文本描述。为此,我们设计了一种分层字幕生成流程,以捕捉传感器数据中的统计、结构与语义信息。该方法促成迄今最大的传感-语言数据集构建,包含超过5970万小时数据,来自逾10.3万名用户。此外,SensorLM扩展了主流多模态预训练架构(如CLIP、CoCa),并将其作为通用架构中的特定变体恢复。在真实世界的人体活动分析与医疗任务中进行的大量实验表明,SensorLM在零样本识别、少样本学习与跨模态检索方面均优于现有最佳方法。该模型还展现出显著的缩放特性、标签高效性、传感器字幕生成能力以及对未见任务的零样本泛化性能。

原文摘要 · Abstract (English)

We present SensorLM, a family of sensor-language foundation models that enable wearable sensor data understanding with natural language. Despite its pervasive nature, aligning and interpreting sensor data with language remains challenging due to the lack of paired, richly annotated sensor-text descriptions in uncurated, real-world wearable data. We introduce a hierarchical caption generation pipeline designed to capture statistical, structural, and semantic information from sensor data. This approach enabled the curation of the largest sensor-language dataset to date, comprising over 59.7 million hours of data from more than 103,000 people. Furthermore, SensorLM extends prominent multimodal pretraining architectures (e.g., CLIP, CoCa) and recovers them as specific variants within a generic architecture. Extensive experiments on real-world tasks in human activity analysis and healthcare verify the superior performance of SensorLM over state-of-the-art in zero-shot recognition, few-shot learning, and cross-modal retrieval. SensorLM also demonstrates intriguing capabilities including scaling behaviors, label efficiency, sensor captioning, and zero-shot generalization to unseen tasks.

可穿戴设备多模态语言模型健康监测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。