arXiv:2511.00301cs.LGphysics.med-ph2025-11被引 2

对比八种不确定性量化方法,提升可穿戴心率信号分析的可信度。

A systematic evaluation of uncertainty quantification techniques in deep learning: a case study in photoplethysmography signal analysis

  • 采用八种不确定性量化技术,针对房颤检测与血压回归任务评估。
  • 局部校准和自适应性评估比全局指标更真实反映模型表现。
  • 适合临床部署场景,尤其关注小样本患者的小尺度可靠性。

深度学习模型在医疗时间序列数据(如可穿戴光电容积脉搏波描记法,PPG)上训练后,可在临床外持续监测生理参数。但实际应用中性能下降可能带来不良患者结果。可靠的预测不确定性可帮助临床医生判断模型输出可信度。为此,本文在两个临床相关任务上——房颤检测(分类)与两种血压回归变体——实施了前所未有的八种不确定性量化(UQ)技术。通过建立全面评估流程,对这些方法进行严格比较。结果显示,不同技术的不确定性可靠性呈现复杂图景,最优方法取决于不确定性表达方式、评估指标及可靠性尺度。我们发现,局部校准和自适应性评估能提供比常用全局可靠性指标更实用的模型行为洞察。强调评估标准应匹配实际应用场景:由于每位患者仅少量测量,需优先保证所选不确定性表达的小规模可靠性,同时尽可能保留预测性能。

原文摘要 · Abstract (English)

In principle, deep learning models trained on medical time-series, including wearable photoplethysmography (PPG) sensor data, can provide a means to continuously monitor physiological parameters outside of clinical settings. However, there is considerable risk of poor performance when deployed in practical measurement scenarios leading to negative patient outcomes. Reliable uncertainties accompanying predictions can provide guidance to clinicians in their interpretation of the trustworthiness of model outputs. It is therefore of interest to compare the effectiveness of different approaches. Here we implement an unprecedented set of eight uncertainty quantification (UQ) techniques to models trained on two clinically relevant prediction tasks: Atrial Fibrillation (AF) detection (classification), and two variants of blood pressure regression. We formulate a comprehensive evaluation procedure to enable a rigorous comparison of these approaches. We observe a complex picture of uncertainty reliability across the different techniques, where the most optimal for a given task depends on the chosen expression of uncertainty, evaluation metric, and scale of reliability assessed. We find that assessing local calibration and adaptivity provides practically relevant insights about model behaviour that otherwise cannot be acquired using more commonly implemented global reliability metrics. We emphasise that criteria for evaluating UQ techniques should cater to the model's practical use case, where the use of a small number of measurements per patient places a premium on achieving small-scale reliability for the chosen expression of uncertainty, while preserving as much predictive performance as possible.

不确定性量化医学信号深度学习可穿戴设备

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。