用小数据模型结合外部生物库,发现未知疾病关联
Improving Diseases Predictions Utilizing External Bio-Banks
- 用10K样本训练模型填补代谢组数据,再用于英国生物银行分析
- 发现吸烟与血管性痴呆的新关联,且模型未接触该信息
- 无需直接疾病标签,即可推断代谢物与肥胖的潜在风险
机器学习在医学等关键领域已取得成功,但常受限于缺乏疾病标签。本研究通过在10K样本上从头训练LightGBM模型,填补代谢组学特征,并应用于英国生物银行(UKBB)进行下游分析。利用填补后的代谢组数据开展生存分析,成功识别出预测模型未察觉的生物学相关性。进一步对关键代谢物进行全基因组关联分析(GWAS),揭示了吸烟与血管性痴呆之间的新关联——尽管此关系为流行病学公认,但未出现在模型训练数据中,验证了方法提取真实信号的能力。同时,将生存模型嵌入10K数据,发现代谢物质与肥胖间的关联,表明即使无直接结局标签,也能推断未来患者疾病风险。结果表明,在数据有限情况下,结合外部生物库、生存分析与遗传研究,仍可挖掘有价值生物洞察。
原文摘要 · Abstract (English)
Machine learning has been successfully used in critical domains, such as medicine. However, extracting meaningful insights from biomedical data is often constrained by the lack of their available disease labels. In this research, we demonstrate how machine learning can be leveraged to enhance explainability and uncover biologically meaningful associations, even when predictive improvements in disease modeling are limited. We train LightGBM models from scratch on our dataset (10K) to impute metabolomics features and apply them to the UK Biobank (UKBB) for downstream analysis. The imputed metabolomics features are then used in survival analysis to assess their impact on disease-related risk factors. As a result, our approach successfully identified biologically relevant connections that were not previously known to the predictive models. Additionally, we applied a genome-wide association study (GWAS) on key metabolomics features, revealing a link between vascular dementia and smoking. Although being a well-established epidemiological relationship, this link was not embedded in the model's training data, which validated the method's ability to extract meaningful signals. Furthermore, by integrating survival models as inputs in the 10K data, we uncovered associations between metabolic substances and obesity, demonstrating the ability to infer disease risk for future patients without requiring direct outcome labels. These findings highlight the potential of leveraging external bio-banks to extract valuable biomedical insights, even in data-limited scenarios. Our results demonstrate that machine learning models trained on smaller datasets can still be used to uncover real biological associations when carefully integrated with survival analysis and genetic studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。