arXiv:2603.26721eess.SPcs.AI2026-03被引 1

用视觉变压器分析心电图,有效缓解个体差异带来的分类难题。

Stress Classification from ECG Signals Using Vision Transformer

  • 将心电图转为二维谱图输入视觉变压器,利用注意力机制捕捉复杂特征。
  • 在WESAD和RML数据集上,三分类准确率达71.01%和76.7%,二分类达88.3%。
  • 端到端训练,无需人工特征,适合生理信号的高鲁棒性建模场景。

视觉变压器在计算机视觉中表现卓越,但尚未被用于心电图(ECG)等生理信号的应激评估。本文将原始ECG数据通过短时傅里叶变换(STFT)转换为二维谱图,并将其划分为图像块输入变压器编码器。同时对比了1D CNN和ResNet-18模型。在WESAD与瑞莫森多媒体实验室(RML)数据集上采用留一人交叉验证(LOSOCV)进行实验。针对个体间差异这一挑战,本研究证明基于2D谱图与变压器注意力机制的方法能显著提升性能。结果表明,该方法在处理个体差异方面优于基于CNN的模型,并以明显优势超越所有先前的最先进方法。所提方法为端到端设计,无需手工特征,可学习鲁棒表示。在三分类任务中,对RML和WESAD数据集分别取得71.01%和76.7%的准确率;在二分类任务中,于WESAD数据集上达到88.3%准确率。

原文摘要 · Abstract (English)

Vision Transformers have shown tremendous success in numerous computer vision applications; however, they have not been exploited for stress assessment using physiological signals such as Electrocardiogram (ECG). In order to get the maximum benefit from the vision transformer for multilevel stress assessment, in this paper, we transform the raw ECG data into 2D spectrograms using short time Fourier transform (STFT). These spectrograms are divided into patches for feeding to the transformer encoder. We also perform experiments with 1D CNN and ResNet-18 (CNN model). We perform leave-onesubject-out cross validation (LOSOCV) experiments on WESAD and Ryerson Multimedia Lab (RML) dataset. One of the biggest challenges of LOSOCV based experiments is to tackle the problem of intersubject variability. In this research, we address the issue of intersubject variability and show our success using 2D spectrograms and the attention mechanism of transformer. Experiments show that vision transformer handles the effect of intersubject variability much better than CNN-based models and beats all previous state-of-the-art methods by a considerable margin. Moreover, our method is end-to-end, does not require handcrafted features, and can learn robust representations. The proposed method achieved 71.01% and 76.7% accuracies with RML dataset and WESAD dataset respectively for three class classification and 88.3% for binary classification on WESAD.

心电图视觉变压器应激识别谱图分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。