arXiv:2602.15967cs.CV2026-02

用自监督学习实现儿科重症监护室无接触心率监测,抗干扰能力强。

Non-Contact Physiological Monitoring in Pediatric Intensive Care Units via Adaptive Masking and Self-Supervised Learning

  • 基于渐进式课程策略与自适应掩码,动态提升重建难度并保留生理相关性。
  • 在500名患儿数据上达到3.2 bpm的平均绝对误差,比标准方法降低42%。
  • 无需人工标注兴趣区域,对遮挡和噪声有强鲁棒性,适合临床部署。

儿科重症监护室(PICUs)持续监测生命体征对早期发现病情恶化至关重要。然而,传统接触式传感器易引发皮肤刺激、感染风险和患者不适。远程光电容积脉搏波(rPPG)可通过面部视频无接触监测心率,但受限于运动伪影、遮挡、光照变化及实验室与临床数据域偏移,在PICUs中应用有限。本文提出一种针对PICU场景的自监督预训练框架,采用渐进式课程策略。方法基于VisionMamba架构,引入轻量级Mamba控制器,生成时空重要性评分,指导概率性图像块采样,动态增加重建难度同时保持生理相关性。为解决临床标注数据不足,采用教师-学生蒸馏机制:在公开数据集上训练的监督专家模型向学生模型提供潜在生理引导。课程分三阶段:干净公开视频、合成遮挡场景、500名患儿的未标注临床视频。该框架相较标准掩码自编码器降低42%平均绝对误差,优于PhysFormer 31%,最终MAE达3.2 bpm。模型无需显式感兴趣区域提取,始终聚焦脉搏丰富区域,且在临床遮挡和噪声下表现稳健。

原文摘要 · Abstract (English)

Continuous monitoring of vital signs in Pediatric Intensive Care Units (PICUs) is essential for early detection of clinical deterioration and effective clinical decision-making. However, contact-based sensors such as pulse oximeters may cause skin irritation, increase infection risk, and lead to patient discomfort. Remote photoplethysmography (rPPG) offers a contactless alternative to monitor heart rate using facial video, but remains underutilized in PICUs due to motion artifacts, occlusions, variable lighting, and domain shifts between laboratory and clinical data. We introduce a self-supervised pretraining framework for rPPG estimation in the PICU setting, based on a progressive curriculum strategy. The approach leverages the VisionMamba architecture and integrates an adaptive masking mechanism, where a lightweight Mamba-based controller assigns spatiotemporal importance scores to guide probabilistic patch sampling. This strategy dynamically increases reconstruction difficulty while preserving physiological relevance. To address the lack of labeled clinical data, we adopt a teacher-student distillation setup. A supervised expert model, trained on public datasets, provides latent physiological guidance to the student. The curriculum progresses through three stages: clean public videos, synthetic occlusion scenarios, and unlabeled videos from 500 pediatric patients. Our framework achieves a 42% reduction in mean absolute error relative to standard masked autoencoders and outperforms PhysFormer by 31%, reaching a final MAE of 3.2 bpm. Without explicit region-of-interest extraction, the model consistently attends to pulse-rich areas and demonstrates robustness under clinical occlusions and noise.

无接触监测rPPG自监督学习儿科医疗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。