用大模型提升远距离生理信号测量的稳定性与准确性
PhysLLM: Harnessing Large Language Models for Cross-Modal Remote Physiological Sensing
- 将大模型与生理信号模块协同优化,通过语义空间对齐实现跨模态融合
- 在四个数据集上达到顶尖性能,显著提升光照变化和运动干扰下的鲁棒性
- 适合做智能健康监测、非接触式生命体征分析的研究者参考
远程光电容积脉搏波描记(rPPG)可实现非接触式生理参数测量,但易受光照变化、运动伪影及时间建模能力不足的影响。大语言模型(LLM)擅长捕捉长程依赖,却因文本导向设计难以处理连续、敏感的rPPG信号。为此,我们提出PhysLLM,一种协同优化框架,融合LLM与领域特定rPPG组件。首先,提出文本原型引导(TPG)策略,将血流动力学特征映射至LLM可理解的语义空间,建立跨模态对齐。其次,设计新型双域平稳(DDS)算法,通过自适应时频域特征重加权解决信号不稳定性问题。最后,系统注入任务特异性生理先验,包括生理统计、环境上下文问答与任务描述,利用跨模态学习融合视觉与文本信息,实现对变光、运动等复杂场景的动态适应。在四个基准数据集上的评估表明,PhysLLM达到当前最优精度与鲁棒性,展现出优异的泛化能力。代码已开源:https://github.com/Alex036225/PhysLLM。
原文摘要 · Abstract (English)
Remote photoplethysmography (rPPG) enables non-contact physiological measurement but remains highly susceptible to illumination changes, motion artifacts, and limited temporal modeling. Large Language Models (LLMs) excel at capturing long-range dependencies, offering a potential solution but struggle with the continuous, noise-sensitive nature of rPPG signals due to their text-centric design. To bridge this gap, we introduce the PhysLLM, a collaborative optimization framework that synergizes LLMs with domain-specific rPPG components. Specifically, the Text Prototype Guidance (TPG) strategy is proposed to establish cross-modal alignment by projecting hemodynamic features into LLM-interpretable semantic space, effectively bridging the representational gap between physiological signals and linguistic tokens. Besides, a novel Dual-Domain Stationary (DDS) Algorithm is proposed for resolving signal instability through adaptive time-frequency domain feature re-weighting. Finally, rPPG task-specific cues systematically inject physiological priors through physiological statistics, environmental contextual answering, and task description, leveraging cross-modal learning to integrate both visual and textual information, enabling dynamic adaptation to challenging scenarios like variable illumination and subject movements. Evaluation on four benchmark datasets, PhysLLM achieves state-of-the-art accuracy and robustness, demonstrating superior generalization across lighting variations and motion scenarios. The source code is available at https://github.com/Alex036225/PhysLLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。