arXiv:2510.10366cs.CVcs.LG2025-10

视觉大模型可高效分析心率信号,实现血压等生命体征的高精度预测。

Vision4PPG: Emergent PPG Analysis Capability of Vision Foundation Models for Vital Signs like Blood Pressure

  • 将一维脉搏波信号转为二维图像,用视觉大模型直接处理。
  • 在血压估计等任务上达到当前最优性能,多个指标超越时序模型。
  • 适用于多种信号可视化形式,适合临床科研人员快速部署使用。

可穿戴与临床设备中的光电容积脉搏波(PPG)传感器能非侵入、实时提供生理信息。以往研究多采用专用或时序类基础模型(FM)进行基准测试。我们通过微调发现,视觉基础模型(VFM)同样适用于此类任务,甚至在多项任务中取得显著优于现有方法的性能,尤其在血压估计方面表现突出。本研究通过将一维PPG信号转换为类似图像的二维表示(如短时傅里叶变换,STFT),利用最新视觉模型(如DINOv3和SIGLIP-2),在多个生命体征及血液检测任务上均获得良好效果。提出的Vision4PPG方法首次系统验证了视觉模型在PPG分析中的通用能力,涵盖六项额外任务,并对比了先进时序模型,证明其在不同二维输入表示(如STFT相位、递归图)下具有强泛化性。借助参数高效微调(PEFT)技术,该方法兼具高性能与计算效率,为临床科学家提供了一套强大且易用的新工具。

原文摘要 · Abstract (English)

Photoplethysmography (PPG) sensor in wearable and clinical devices provides valuable physiological insights in a non-invasive and real-time fashion. Specialized Foundation Models (FM) or repurposed time-series FMs are used to benchmark physiological tasks. Our experiments with fine-tuning FMs reveal that Vision FM (VFM) can also be utilized for this purpose and, in fact, surprisingly leads to state-of-the-art (SOTA) performance on many tasks, notably blood pressure estimation. We leverage VFMs by simply transforming one-dimensional PPG signals into image-like two-dimensional representations, such as the Short-Time Fourier transform (STFT). Using the latest VFMs, such as DINOv3 and SIGLIP-2, we achieve promising performance on other vital signs and blood lab measurement tasks as well. Our proposal, Vision4PPG, unlocks a new class of FMs to achieve SOTA performance with notable generalization to other 2D input representations, including STFT phase and recurrence plots. Our work improves upon prior investigations of vision models for PPG by conducting a comprehensive study, comparing them to state-of-the-art time-series FMs, and demonstrating the general PPG processing ability by reporting results on six additional tasks. Thus, we provide clinician-scientists with a new set of powerful tools that is also computationally efficient, thanks to Parameter-Efficient Fine-Tuning (PEFT) techniques.

视觉模型生命体征脉搏波大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。