用隐空间编辑视频中生理信号,保护隐私同时保持画面质量。
Editing Physiological Signals in Videos Using Latent Representations
- 通过3D VAE编码视频,结合目标心率提示进行可控编辑。
- 重建后视频心率误差仅10.00 bpm,PSNR达38.96 dB,视觉保真度高。
- 适合用于生物特征匿名化或生成带指定生命体征的视频。
基于摄像头的生理信号估计为无接触监测心率(HR)提供了便利手段。然而,面部视频中包含的生命体征会引发严重隐私问题,可能暴露个体健康与情绪状态等敏感信息。为此,我们提出一种学习框架,在保留视觉保真度的前提下编辑视频中的生理信号。首先,利用预训练的3D变分自编码器(3D VAE)将输入视频编码至隐空间,同时通过冻结的文本编码器嵌入目标心率提示。采用可训练的时空层与自适应层归一化(AdaLN)融合二者,以捕捉远程光电容积脉搏波图(rPPG)信号的强时序一致性。在解码器中引入特征逐元素线性调制(FiLM)及微调输出层,避免重建过程中的生理信号退化,实现精确的生理调节。实验表明,本方法在选定数据集上平均保持38.96 dB PSNR与0.98 SSIM,使用先进rPPG估计算法时,心率调控平均绝对误差(MAE)为10.00 bpm,平均绝对百分比误差(MAPE)为10.09%。该设计支持可控心率编辑,适用于真实视频中生物特征匿名化或合成具有特定生命体征的逼真视频。
原文摘要 · Abstract (English)
Camera-based physiological signal estimation provides a non-contact and convenient means to monitor Heart Rate (HR). However, the presence of vital signals in facial videos raises significant privacy concerns, as they can reveal sensitive personal information related to the health and emotional states of an individual. To address this, we propose a learned framework that edits physiological signals in videos while preserving visual fidelity. First, we encode an input video into a latent space via a pretrained 3D Variational Autoencoder (3D VAE), while a target HR prompt is embedded through a frozen text encoder. We fuse them using a set of trainable spatio-temporal layers with Adaptive Layer Normalizations (AdaLN) to capture the strong temporal coherence of remote Photoplethysmography (rPPG) signals. We apply Feature-wise Linear Modulation (FiLM) in the decoder with a fine-tuned output layer to avoid the degradation of physiological signals during reconstruction, enabling accurate physiological modulation in the reconstructed video. Empirical results show that our method preserves visual quality with an average PSNR of 38.96 dB and SSIM of 0.98 on selected datasets, while achieving an average HR modulation error of 10.00 bpm MAE and 10.09% MAPE using a state-of-the-art rPPG estimator. Our design's controllable HR editing is useful for applications such as anonymizing biometric signals in real videos or synthesizing realistic videos with desired vital signs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。