arXiv:2509.00221cs.LGeess.AS2025-09被引 1

用语音模型处理可穿戴设备数据,效果超越专用模型。

Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data

  • 复用语音模型的特征编码器提取传感器时序特征
  • 在情绪分类等任务上达到当前最佳性能
  • 特别适合小样本时序数据分析场景

语音与可穿戴传感器时序数据均在时域和频域蕴含信息,如谱功率和波形片段。我们发现语音基础模型学习的表征可泛化至非语音领域,在多种可穿戴传感器时序任务中表现卓越。使用HuBERT和wav2vec 2.0提取的特征,其上的探测器在情绪分类、心律失常检测和活动分类任务中优于直接在特定模态数据上训练的自监督模型。研究发现,语音模型中的卷积特征编码器对可穿戴传感器应用尤为关键。该方法通过简单的探测机制,显著提升数据稀缺时序任务的表现。本工作为统一语音与传感器模态的通用时序建模迈出了重要一步。

原文摘要 · Abstract (English)

Both speech and sensor time series data encode information in both the time- and frequency- domains, like spectral powers and waveform shapelets. We show that speech foundation models learn representations that generalize beyond the speech domain and achieve state-of-the-art performance on diverse time-series tasks from wearable sensors. Probes trained on features extracted from HuBERT and wav2vec 2.0 outperform those extracted from self-supervised models trained directly on modality-specific datasets for mood classification, arrhythmia detection, and activity classification tasks. We find that the convolutional feature encoders of speech models are particularly relevant for wearable sensor applications. The proposed approach enhances performance on data-scarce time-series tasks using simple probing methods. This work takes a step toward developing generalized time-series models that unify speech and sensor modalities.

时序建模语音模型可穿戴设备迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。