对比六种深度学习模型,为可穿戴设备行为预测选型提供实证指导。
A Comparative Study of Deep Learning Architectures for Multi-Horizon Behavioural Forecasting for Mobile Health
- 在三数据集上对比六种模型,评估多天预测表现
- 零样本模型TimesFM表现接近训练模型,低数据时优势明显
- 个体微调可降16%-60%误差,睡眠预测受益最显著
可穿戴设备和智能手机生成丰富的行为时间序列,可用于主动健康干预,但针对这些数据的现代预测架构系统性比较仍不足。本文在涵盖800多名参与者、三个公开数据集上,对六种深度学习架构、两种零样本基础模型(FM)及统计基线进行评测,报告步数、屏幕使用时间和睡眠时长在1-8天预测窗口下的各特征指标。进一步开展所有模型的个体化微调研究,并评估基础模型在不同数据规模和时间粒度下的迁移能力。关键发现:(i) 无单一模型全面领先,训练模型中PatchTST最优,其余三个(TCN、MLP、Transformer)性能无显著差异;(ii) 零样本模型TimesFM在零样本下表现与训练模型相当,尤其在低数据场景更优;(iii) 个体级微调使各特征均方根误差降低16%-60%,睡眠预测改善最大,步数预测最小。结果为移动健康预测中的模型选择、基础模型应用及个性化策略提供实践依据。据我们所知,这是首个联合评估现代深度学习、基础模型与个性化策略在多时段行为预测中表现的研究。
原文摘要 · Abstract (English)
Wearable devices and smartphones generate rich behavioural time series that can support proactive health interventions, yet systematic comparisons of modern forecasting architectures for these data are lacking. In particular, it remains unclear how models generalise across populations, how different architectures respond to participant-level fine-tuning and how forecasting accuracy degrades across multi-day horizons. We benchmark six deep learning architectures, two zero-shot Foundation Models (FM) and statistical baselines on three public datasets encompassing over 800 participants, reporting per-feature metrics for step counts, screen time and sleep duration across 1-8 day horizons. We further conduct a per-feature personalisation study across all six architectures and assess FM transferability across dataset sizes and temporal granularities. Our key findings are: (i) no single architecture dominates, PatchTST leads among trained models while the three runners-up (TCN, MLP, Transformer) show no meaningful performance difference; (ii) the FM TimesFM matches or exceeds trained models zero-shot, especially in low-data regimes and (iii) participant-level fine-tuning reduces per-feature RMSE by 16-60\%, with sleep benefiting most and step counts least. These results provide practical guidance on architecture selection, FM applicability and personalisation strategies for mobile health forecasting. To the best of our knowledge, this is the first study to jointly evaluate modern deep learning, FMs and personalisation for multi-horizon behavioural forecasting from wearables.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。