对比通用与专用时序模型在生理信号中的表现,发现专用模型更优。
Generalist vs Specialist Time Series Foundation Models: Investigating Potential Emergent Behaviors in Assessing Human Health Using PPG Signals
- 用多领域数据训练通用模型,与专用模型对比性能。
- 专用模型在51项任务中胜出27%的评估指标得分。
- 适合关注健康监测、时序建模的科研与工程人员参考。
基础模型是大规模预训练的机器学习模型,可适应多种下游任务。在自然语言处理和计算机视觉中已有广泛应用,如GPT、BERT和CLIP。如今,它们也日益被引入时序分析,尤其在生理传感领域。然而,多数时序基础模型为专用模型——预训练与测试数据同类型,如心电图、脑电图和光电容积脉搏波(PPG)。近期工作如MOMENT则使用天气、交通、电力等多领域数据训练通用时序基础模型。本文开展全面基准测试,比较通用与专用模型在PPG信号上的表现。通过覆盖心脏状态评估、实验室值估计和跨模态推理的共51项任务,从胜率、平均性能、特征质量、微调增益、性能方差、可迁移性与可扩展性七个维度进行评估。这些指标综合反映模型能力、适应性、鲁棒性与效率。全微调场景下,专用模型胜率高出27%。进一步分析了泛化能力、公平性、注意力可视化及训练数据选择的重要性。
原文摘要 · Abstract (English)
Foundation models are large-scale machine learning models that are pre-trained on massive amounts of data and can be adapted for various downstream tasks. They have been extensively applied to tasks in Natural Language Processing and Computer Vision with models such as GPT, BERT, and CLIP. They are now also increasingly gaining attention in time-series analysis, particularly for physiological sensing. However, most time series foundation models are specialist models - with data in pre-training and testing of the same type, such as Electrocardiogram, Electroencephalogram, and Photoplethysmogram (PPG). Recent works, such as MOMENT, train a generalist time series foundation model with data from multiple domains, such as weather, traffic, and electricity. This paper aims to conduct a comprehensive benchmarking study to compare the performance of generalist and specialist models, with a focus on PPG signals. Through an extensive suite of total 51 tasks covering cardiac state assessment, laboratory value estimation, and cross-modal inference, we comprehensively evaluate both models across seven dimensions, including win score, average performance, feature quality, tuning gain, performance variance, transferability, and scalability. These metrics jointly capture not only the models' capability but also their adaptability, robustness, and efficiency under different fine-tuning strategies, providing a holistic understanding of their strengths and limitations for diverse downstream scenarios. In a full-tuning scenario, we demonstrate that the specialist model achieves a 27% higher win score. Finally, we provide further analysis on generalization, fairness, attention visualizations, and the importance of training data choice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。