自预训练能提升医疗时间序列模型准确率,尤其在数据少时效果更明显。
Is Self-Pretraining really useful to improve diagnosis in medical Time Series?

- 用掩码策略在医疗时间序列上做自预训练,增强时序与跨模态表征
- 在3个数据集上准确率提升0-6个百分点,深度模型受益更显著
- 无需修改架构,适合临床小样本场景的模型优化
受长上下文基准中Transformer自预训练(SPT)有效性的启发,本文研究SPT是否适用于多模态、多变量甚至简单单变量医疗时间序列。目标是评估SPT对不同医疗任务中基于Transformer模型性能与可扩展性的影响,尤其在数据有限条件下的表现。我们在三个代表性任务上进行评估:康复机器人(Camargo数据集)、压力检测(Non-EEG Stress)和帕金森病检测(Gait Parkinson's Disease)。模型从零开始训练或通过四种基于掩码的SPT目标进行预训练,系统性地改变模型深度以分析容量与预训练收益的交互关系。在各数据集与配置下,SPT均使分类准确率提升0-6个百分点,且不仅在多变量设置中有效,即使仅使用单变量输入也观察到增益。深层模型因能更好利用预训练中学习到的时序表示而获益更多。结果表明,SPT是一种简单通用的策略,可在不改变任务特定架构的前提下提升医疗时间序列模型性能,支持其在数据稀缺的临床环境中提升鲁棒性与准确性。
原文摘要 · Abstract (English)
Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate whether similar gains extend to multimodal, multivariate, and even simple univariate medical time series. Our objective is to assess the impact of SPT on the performance and scalability of transformer-based models across diverse medical applications, particularly under limited data conditions. We evaluate transformer architectures on three representative medical time-series tasks: rehabilitation robotics (Camargo dataset), stress detection (Non-EEG Stress), and Parkinson's disease detection (Gait Parkinson's Disease). Models are trained either from scratch or through SPT using four masking-based objectives designed to promote temporal and cross-modal representation learning, and we systematically vary model depth to examine how capacity interacts with pre-training benefits. Across datasets and configurations, SPT consistently improves classification accuracy by 0-6 percentage points depending on masking strategy, dataset and architecture, with gains observed not only in multivariate settings but also when models are restricted to simple univariate inputs. The improvements increase for deeper models that can better exploit the enriched temporal representations learned during pre-training. These findings indicate that SPT is a simple and general strategy that enhances transformer performance on medical time-series tasks without requiring task-specific architectural changes, supporting its potential to improve robustness and accuracy in data-limited clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。