融合3D CNN与Transformer,提升视频远程心率测量精度
VidFormer: A novel end-to-end framework fused by 3DCNN and Transformer for Video-based Remote Physiological Measurement
- 结合3D CNN提取局部时空特征,Transformer捕捉全局依赖关系
- 在5个公开数据集上均超越现有最优方法,平均相关系数达0.942
- 对肤色、妆容、运动等实际干扰因素鲁棒,适合真实场景应用
基于面部视频的远程生理信号测量(rPPG)旨在从视频中推断面部血流变化。尽管现有深度学习方法表现良好,但受限于卷积神经网络(CNN)和Transformer的固有缺陷,在小规模与大规模数据集间难以兼顾性能。本文提出VidFormer,一种融合3D卷积神经网络(3DCNN)与Transformer的端到端框架。首先分析传统皮肤反射模型,并提出改进版本以重建rPPG信号。基于此,VidFormer分别利用3DCNN和Transformer提取输入数据的局部与全局特征。为增强时空特征提取能力,设计了针对3DCNN和Transformer的时空间注意力机制,并引入模块实现两者间的信息交换与融合。在五个公开数据集上的评估显示,VidFormer显著优于当前最优方法。最后,分析各模块作用,并考察肤色、妆容及运动对性能的影响。
原文摘要 · Abstract (English)
Remote physiological signal measurement based on facial videos, also known as remote photoplethysmography (rPPG), involves predicting changes in facial vascular blood flow from facial videos. While most deep learning-based methods have achieved good results, they often struggle to balance performance across small and large-scale datasets due to the inherent limitations of convolutional neural networks (CNNs) and Transformer. In this paper, we introduce VidFormer, a novel end-to-end framework that integrates 3-Dimension Convolutional Neural Network (3DCNN) and Transformer models for rPPG tasks. Initially, we conduct an analysis of the traditional skin reflection model and subsequently introduce an enhanced model for the reconstruction of rPPG signals. Based on this improved model, VidFormer utilizes 3DCNN and Transformer to extract local and global features from input data, respectively. To enhance the spatiotemporal feature extraction capabilities of VidFormer, we incorporate temporal-spatial attention mechanisms tailored for both 3DCNN and Transformer. Additionally, we design a module to facilitate information exchange and fusion between the 3DCNN and Transformer. Our evaluation on five publicly available datasets demonstrates that VidFormer outperforms current state-of-the-art (SOTA) methods. Finally, we discuss the essential roles of each VidFormer module and examine the effects of ethnicity, makeup, and exercise on its performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。