arXiv:2608.29663cs.CV2026-08

用视觉语言模型增强脉搏信号,抗干扰能力更强。

PhysVR: Vision-Language Model Guided Interference-aware Temporal Feature Refinement for Remote Physiological Measurement

论文配图:PhysVR: Vision-Language Model Guided Interference-aware Temporal Feature Refinement for Remote Physiological Measurement
图 1 · 摘自论文原文
  • 用视觉语言模型提取干扰证据,动态优化特征
  • 跨数据集测试下精度提升,最高达7.2%相对改进
  • 适合需要高鲁棒性的远程生理监测场景

远程光电容积脉搏波描记(rPPG)可通过面部视频实现无接触生理测量,但其微弱的脉搏相关变化易受光照变化、头部运动、面部模糊和感兴趣区域不稳等因素干扰。现有方法多在特征学习阶段抑制干扰,却未关注所学时序特征是否仍受干扰影响,以及如何在rPPG估计前进一步抑制干扰。为此,本文提出PhysVR:一种基于视觉-语言模型引导的干扰感知时序特征精炼框架。具体而言,生理主干网络生成全局时序特征与粗略的rPPG预测,从中构建基于局部时序特性的生理可靠性证据;同时,冻结的视觉-语言模型在干扰导向提示下处理采样的面部帧,通过证据头从其输出中提取视觉干扰证据。时序交叉注意力将生理与视觉证据融合至全局时序特征,构建干扰感知的时序上下文。该上下文指导一个共享的时序修正单元进行通用精炼,同时四个干扰特异性专家通过自适应路由选择性抑制不同干扰。最终的精炼时序特征用于精确的rPPG估计。在五个公开基准上的大量实验表明,PhysVR在跨数据集和同数据集评估协议下均持续优于代表性方法。

原文摘要 · Abstract (English)

Remote photoplethysmography (rPPG) enables contactless physiological measurement from facial videos, yet its subtle pulse-related variations are easily affected by illumination variation, head motion, facial blur, and region-of-interest instability. Existing methods mainly suppress interference during feature learning, while whether the learned temporal features remain affected by interference and how to further suppress such interference before rPPG estimation are rarely examined. To address this limitation, we propose PhysVR, a vision-language model guided interference-aware temporal feature refinement framework for rPPG estimation. Specifically, a physiological backbone produces global temporal features and a coarse rPPG prediction, from which signal-derived physiological reliability evidence is constructed from local temporal characteristics. In parallel, a frozen vision-language model processes sampled facial frames under an interference-oriented prompt, and an evidence head extracts visual interference evidence from the VLM output. Temporal cross-attention integrates the physiological and visual evidence with the global temporal features to construct interference-aware temporal context. Guided by this context, a shared temporal correction unit performs general refinement, while four interference-specific experts selectively suppress different interference through adaptive routing. The refined temporal features are then used for final rPPG estimation. Extensive experiments on five public benchmarks demonstrate that PhysVR consistently outperforms representative methods under both intra-dataset and cross-dataset evaluation protocols.

rPPG视觉语言模型时序精炼生理监测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。