首次建立可穿戴动作识别的缩放定律,揭示数据多样性比单用户数据量更重要。
Scaling laws in wearable human activity recognition
- 通过大规模网格搜索,发现预训练损失与数据量、参数量呈幂律关系。
- 增加用户数量带来的性能提升优于增加单个用户数据量。
- 适用于研究可穿戴设备动作识别的模型设计与实验复现。
针对可穿戴多模态传感器的人体动作识别(HAR),已有大量深度架构和自监督预训练方法被提出。缩放定律有望通过关联模型容量与预训练数据量,推动更系统化的设计。然而,目前在语言和视觉领域已建立的缩放定律尚未在HAR中充分验证。本文通过在预训练数据量和Transformer架构上进行全面网格搜索,首次确立了HAR领域的缩放定律。结果表明,预训练损失与数据量及参数量呈幂律关系;增加数据集中的用户数量所带来的性能提升,显著高于仅增加每个用户的数据量,说明预训练数据的多样性至关重要,这与部分先前自监督HAR研究的结论相悖。这些缩放定律在三个基准数据集(UCI HAR、WISDM Phone 和 WISDM Watch)上实现了下游任务性能的提升。最后,我们建议基于这些规律重新审视部分既有研究,以采用更充分的模型容量。
原文摘要 · Abstract (English)
Many deep architectures and self-supervised pre-training techniques have been proposed for human activity recognition (HAR) from wearable multimodal sensors. Scaling laws have the potential to help move towards more principled design by linking model capacity with pre-training data volume. Yet, scaling laws have not been established for HAR to the same extent as in language and vision. By conducting an exhaustive grid search on both amount of pre-training data and Transformer architectures, we establish the first known scaling laws for HAR. We show that pre-training loss scales with a power law relationship to amount of data and parameter count and that increasing the number of users in a dataset results in a steeper improvement in performance than increasing data per user, indicating that diversity of pre-training data is important, which contrasts to some previously reported findings in self-supervised HAR. We show that these scaling laws translate to downstream performance improvements on three HAR benchmark datasets of postures, modes of locomotion and activities of daily living: UCI HAR and WISDM Phone and WISDM Watch. Finally, we suggest some previously published works should be revisited in light of these scaling laws with more adequate model capacities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。