用新框架筛选高质量视频,让无监督心率检测更准
rPPG-VQA: A Video Quality Assessment Framework for Unsupervised rPPG Training
- 分信号与场景双路评估视频质量,结合多方法共识和大模型识别干扰
- 在真实视频上训练后,标准测试集准确率显著提升
- 适合做无监督rPPG研究的团队使用,尤其关注数据预处理
无监督远程光电容积脉搏波(rPPG)可利用未标注视频数据,但其性能受真实场景低质视频严重制约。当前缺乏对视频是否适合作为rPPG训练数据的评估机制。现有视频质量评估方法主要面向人类感知,不适用于此目标。本文提出rPPG-VQA框架,融合信号级与场景级分析,采用双分支架构:信号级分支通过多方法共识机制实现鲁棒信噪比(SNR)估计;场景级分支利用多模态大语言模型(MLLM)识别运动、光照不稳定等干扰。此外,设计两阶段自适应采样(TAS)策略,根据质量评分筛选最优训练集。实验表明,经本框架过滤的大规模真实视频训练后,无监督rPPG模型在标准基准上准确率大幅提升。代码已开源。
原文摘要 · Abstract (English)
Unsupervised remote photoplethysmography (rPPG) promises to leverage unlabeled video data, but its potential is hindered by a critical challenge: training on low-quality "in-the-wild" videos severely degrades model performance. An essential step missing here is to assess the suitability of the videos for rPPG model learning before using them for the task. Existing video quality assessment (VQA) methods are mainly designed for human perception and not directly applicable to the above purpose. In this work, we propose rPPG-VQA, a novel framework for assessing video suitability for rPPG. We integrate signal-level and scene-level analyses and design a dual-branch assessment architecture. The signal-level branch evaluates the physiological signal quality of the videos via robust signal-to-noise ratio (SNR) estimation with a multi-method consensus mechanism, and the scene-level branch uses a multimodal large language model (MLLM) to identify interferences like motion and unstable lighting. Furthermore, we propose a two-stage adaptive sampling (TAS) strategy that utilizes the quality score to curate optimal training datasets. Experiments show that by training on large-scale, "in-the-wild" videos filtered by our framework, we can develop unsupervised rPPG models that achieve a substantial improvement in accuracy on standard benchmarks. Our code is available at https://github.com/Tianyang-Dai/rPPG-VQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。