用手机视频非侵入式测生理指标,精度随数据增多稳步提升
Full-Self Diagnostics (FSD): Physics-Grounded Visual Biomarker Inference from Smartphone Video via Inverse Problems and Operator Learning
- 基于物理模型和逆问题求解,从人脸视频推断生理状态
- 实测血糖预测误差仅17%-29.86%,97.57%结果在安全区
- 无需专业设备,适合糖尿病等慢病的居家长期监测
我们提出全自诊断(FSD)框架,从消费者智能手机拍摄的9秒无约束面部视频中恢复潜在生理状态。该方法整合五部分:(1)基于辐射传输方程和色素吸收的物理前向模型,将相机观测映射到生物标志物浓度;(2)信息论可观测性理论证明多通道视觉信号(光谱、脉搏、呼吸、微表情、眼动)与生理状态互信息逐级增加;(3)带Tikhonov正则化的稳定逆问题求解,具备域内统一可识别性保证;(4)操作符学习形式实现跨设备、分辨率与人群泛化;(5)监督学习过程可解释为随机变分推断,性能随配对观测数平方根倒数提升。在59名受试者共38,812组真实世界配对扫描中验证。主作者自采数据(血糖35-550 mg/dL)显示,平均绝对相对偏差(MARD)为29.86%,97.57%预测位于克拉克误差网格区A+B,仅0.27%进入危险区E。一名管理良好的糖尿病患者在70-180 mg/dL区间实现17% MARD。结果表明,消费级面部视频在完全无约束条件下蕴含足够结构化信息,支持临床相关、非侵入式生物标志物推断,且性能可预测地随数据增加而改善。
原文摘要 · Abstract (English)
We present Full-Self Diagnostics (FSD), a unified mathematical framework for recovering latent physiological states from unconstrained 9-second facial videos captured by consumer smartphones. The approach integrates five mutually reinforcing components: (1) a physics-based forward model derived from the radiative transfer equation and chromophore absorption that maps camera observables to biomarker concentrations; (2) an information-theoretic observability theory proving that multi-channel visual signals (spectral, pulse, respiratory, micro-expression, and oculomotor) contain strictly increasing mutual information with physiological state; (3) a stable, Tikhonov-regularized inverse problem with domain-uniform identifiability guarantees; (4) an operator-learning formulation that enables generalization across devices, resolutions, and populations; and (5) a supervised learning procedure, interpretable as stochastic variational inference, that continuously refines the model from paired biosensor ground truth with performance improving proportionally to one over the square root of the number of paired observations. Empirical validation on 38812 real-world paired scans across 59 subjects demonstrates practical performance. Self-collected data from the lead author (glucose range 35-550 mg/dL) yields MARD of 29.86 percent with 97.57 percent of predictions in Clarke Error Grid Zones A+B and only 0.27 percent in the dangerous Zone E. A well-managed diabetic participant achieves MARD of 17 percent in the narrower 70-180 mg/dL band. These results confirm that consumer-grade facial video encodes sufficient structured information for clinically relevant, non-invasive biomarker inference under fully unconstrained conditions, with performance scaling predictably as more paired data becomes available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。