arXiv:2506.17620cs.LGcs.CY2025-06

用生活方式数据预测13种慢病风险,结果经医学文献验证更可信。

Trustworthy Chronic Disease Risk Prediction For Self-Directed Preventive Care via Medical Literature Validation

  • 仅用个人生活习惯数据训练模型,支持自我健康评估
  • 13种慢病预测中关键影响因素与医学文献高度一致
  • 通过文献验证提升模型可信度,适合关注预防的人群

慢性病是长期可管理但通常不可治愈的疾病,亟需有效预防策略。机器学习广泛用于个体慢性病风险评估,但多数依赖血液检查等医疗数据,限制了其在主动自评中的应用。此外,为赢得公众信任,模型需具备可解释性和透明性。尽管部分自评模型包含可解释性,但其解释未经过权威医学文献验证,降低了可靠性。为此,我们构建了基于深度学习的模型,仅使用个人和生活方式因素预测13种慢性病风险,实现可及的自我导向预防护理。关键的是,我们采用基于SHAP的可解释性方法识别最具影响力的特征,并与既有医学文献进行比对。结果显示,模型中最重要特征与医学文献高度一致,且这一现象在13种不同疾病中均成立,表明该方法在慢性病预测中具有广泛可信性。本工作为开发可信的自我预防型机器学习工具奠定基础,未来研究可探索其他可信性增强方法及伦理使用路径。

原文摘要 · Abstract (English)

Chronic diseases are long-term, manageable, yet typically incurable conditions, highlighting the need for effective preventive strategies. Machine learning has been widely used to assess individual risk for chronic diseases. However, many models rely on medical test data (e.g. blood results, glucose levels), which limits their utility for proactive self-assessment. Additionally, to gain public trust, machine learning models should be explainable and transparent. Although some research on self-assessment machine learning models includes explainability, their explanations are not validated against established medical literature, reducing confidence in their reliability. To address these issues, we develop deep learning models that predict the risk of developing 13 chronic diseases using only personal and lifestyle factors, enabling accessible, self-directed preventive care. Importantly, we use SHAP-based explainability to identify the most influential model features and validate them against established medical literature. Our results show a strong alignment between the models' most influential features and established medical literature, reinforcing the models' trustworthiness. Critically, we find that this observation holds across 13 distinct diseases, indicating that this machine learning approach can be broadly trusted for chronic disease prediction. This work lays the foundation for developing trustworthy machine learning tools for self-directed preventive care. Future research can explore other approaches for models' trustworthiness and discuss how the models can be used ethically and responsibly.

慢病预测可解释性自我护理文献验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。