arXiv:2511.02849eess.SPcs.CV2025-11被引 7

用改进数据提升糖尿病低血糖预测准确率,性能比原始数据高2-3%。

Benchmarking ResNet for Short-Term Hypoglycemia Classification with DiaData

  • 通过插值与异常值处理优化15个数据集的糖尿病血糖数据质量
  • 基于清理后数据训练的ResNet模型可提前2小时预测低血糖事件
  • 适合研究糖尿病智能预警与医疗数据清洗的从业者参考

个体化治疗依赖于医疗数据分析,为1型糖尿病(T1D)患者提供情境洞察。然而,异常值、噪声和小样本量难以支撑可靠分析,亟需高质量大规模数据。本研究针对包含2510名T1D患者葡萄糖数据的DiaData数据集进行质量提升:1)采用四分位距法识别异常值并替换为缺失值;2)小于等于25分钟的空缺用线性插值填补,大于等于30分钟且小于120分钟的空缺使用Stineman插值,视觉对比显示其在大间隔下更真实;3)数据清洗后分析葡萄糖与心率相关性,发现低血糖前15至60分钟间存在中等程度关联;4)最终基于主数据库与子数据库II训练状态先进的ResNet模型,实现提前最多2小时的低血糖分类预测,使用更多数据使性能提升7%,使用高质量数据相较原始数据提升2-3%。

原文摘要 · Abstract (English)

Individualized therapy is driven forward by medical data analysis, which provides insight into the patient's context. In particular, for Type 1 Diabetes (T1D), which is an autoimmune disease, relationships between demographics, sensor data, and context can be analyzed. However, outliers, noisy data, and small data volumes cannot provide a reliable analysis. Hence, the research domain requires large volumes of high-quality data. Moreover, missing values can lead to information loss. To address this limitation, this study improves the data quality of DiaData, an integration of 15 separate datasets containing glucose values from 2510 subjects with T1D. Notably, we make the following contributions: 1) Outliers are identified with the interquartile range (IQR) approach and treated by replacing them with missing values. 2) Small gaps ($\le$ 25 min) are imputed with linear interpolation and larger gaps ($\ge$ 30 and $<$ 120 min) with Stineman interpolation. Based on a visual comparison, Stineman interpolation provides more realistic glucose estimates than linear interpolation for larger gaps. 3) After data cleaning, the correlation between glucose and heart rate is analyzed, yielding a moderate relation between 15 and 60 minutes before hypoglycemia ($\le$ 70 mg/dL). 4) Finally, a benchmark for hypoglycemia classification is provided with a state-of-the-art ResNet model. The model is trained with the Maindatabase and Subdatabase II of DiaData to classify hypoglycemia onset up to 2 hours in advance. Training with more data improves performance by 7% while using quality-refined data yields a 2-3% gain compared to raw data.

糖尿病预测数据清洗ResNet

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。