用分层自监督学习提升心电图多变量时序分析效果
Hierarchical Self-Supervised Representation Learning Framework for Multivariate Time Series Grounded in ECG Analysis
- 分两阶段构建时序表征,先局部后全局
- 预训练18万条心电数据,下游任务性能领先
- 结构设计避免特征过度平滑,适合医疗时序分析
医学数据常面临目标数据量少、大量未标注数据分布广泛的问题。自监督学习(SSL)在利用大规模数据方面表现优异,尤其适用于心电图(ECG)分析。本文提出事件重构联合嵌入预测架构(ER-JEPA),一种轻量级多变量时间序列自监督框架,其名称与双层级结构受心脏病学诊断流程启发。该框架包含:(1) 两阶段结构,先对每个时间区间构建表征,再将这些表征作为单变量时间序列处理;(2) 双重联合嵌入预测架构(JEPAs)的层次化集成;(3) 基于视觉变换器(ViT)的主干网络。两个JEPAs的结构串联使模型成为分层联合嵌入预测架构(H-JEPA),旨在编码多层次抽象表示以增强复杂任务的预测能力。本研究成功将H-JEPA应用于12导联心电图数据,作为多变量时间序列建模,并分析了预训练阶段中分层表征的敏感性。此外,定性结果显示,ER-JEPA第一模块生成的中间表征在局部特征提取上表现优异,因其结构上避免了过平滑问题。模型在约18万条10秒记录上预训练,于ST-MEM基准测试中达到最先进下游性能,兼具快速计算与低资源消耗。
原文摘要 · Abstract (English)
Data analysis in the medical domain often encounters scenarios involving a limited target dataset and a large, unannotated dataset with a general distribution. Under such circumstances, self-supervised learning (SSL) methods are highly effective for utilizing large datasets, making them a popular choice for electrocardiogram (ECG) analysis. This work presents the Event Reconstruction Joint-Embedding Predictive Architecture (ER-JEPA), a lightweight SSL framework for multivariate time series, whose name and two-fold hierarchical structure are inspired by the diagnostic approach of cardiologists. At its core, ER-JEPA features: (1) a two-stage structure that constructs representations for each time interval and subsequently processes these representations as a univariate time series, (2) the hierarchical integration of two Joint-Embedding Predictive Architectures (JEPAs), and (3) a Vision Transformer (ViT) backbone. The structural concatenation of two JEPAs categorizes the model as a Hierarchical JEPA (H-JEPA), designed to encode multiple levels of abstract representations for enhanced prediction on complex tasks. This study reports a successful application of H-JEPA to 12-lead ECG data as a multivariate time series, alongside an analysis of the sensitivity of hierarchical representation during the pretraining stage. Furthermore, this study provides a qualitative demonstration that the intermediate representations produced by the first module of ER-JEPA excel at local feature extraction, as they are structurally free from over-smoothing. Pretrained on approximately 180,000 10-second recordings, the model achieves state-of-the-art downstream performance on the ST-MEM benchmark, with rapid computation and minimal resource usage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。