arXiv:2605.18701cs.LGq-bio.QM2026-05

用个体+群体数据生成更准的血检参考值,避免误判。

Learning Normal Representations for Blood Biomarkers

论文配图:Learning Normal Representations for Blood Biomarkers
图 1 · 摘自论文原文
  • 用条件Transformer融合个人病史与人群正常波动数据
  • 个性化参考区间可将68%检测误判为异常,但无临床后果关联
  • 新方法在预测死亡、肾损伤等结果上更精准,适合临床医生使用

血检生物标志物支撑临床诊断,但传统依赖固定人群参考区间,忽视个体稳定变异,易漏诊。现有个性化方法常因数据稀疏导致过拟合,误报率高,并可能包含未察觉疾病。本文利用北美、中东和东亚超160万人群近20亿条纵向检验数据,发现实验室数值高度个体化,纯个性化区间会将高达68%的测量值错误标记为异常,且无对应不良临床结局。为此提出NORMA框架——基于条件Transformer,结合个体历史与人群正常变异信息生成参考区间。该方法显著提升对死亡、急性肾损伤及慢性病的预测精度。研究警示过度个性化风险,表明锚定个体轨迹于人群先验优于单一方法。模型、代码与交互界面已公开,推动透明化个体化检验解读。

原文摘要 · Abstract (English)

Blood-based biomarkers underpin clinical diagnosis and management, yet their interpretation relies largely on fixed population reference intervals that ignore stable, intra-patient variability. As such, population-based interpretation can mask meaningful deviation from an individual's baseline, risking delayed disease detection. To remedy this, there have been increasing efforts to personalize blood biomarker interpretation using individual testing histories. However, these methods may overfit to sparse data, inflating false-positive rates and unnecessary follow-up, and can also unwittingly include unrecognized or subclinical disease. Here, we leverage nearly 2 billion longitudinal laboratory measurements from over 1.6 million individuals across North America, the Middle East, and East Asia, to show that while laboratory values are highly individual, purely personalized intervals routinely overfit, classifying up to 68% of measurements as abnormal, without corresponding associations with adverse clinical outcomes. We then introduce NORMA, a conditional transformer-based framework that generates reference intervals by conditioning on both a patient's history and population-level data about "normal" variation. NORMA-derived intervals achieve higher precision for predicting outcomes, including mortality, acute kidney injury, and chronic disease. These findings caution against over-personalization in laboratory medicine and demonstrate that anchoring individual trajectories to population-level priors outperforms either approach alone. To promote transparency, we publicly release the model, code, and an interactive user interface for accessible, individualized laboratory interpretation.

血检分析个性化医疗时间序列临床预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。