arXiv:2509.07330cs.LGcs.AI2025-09

为年龄性别设计通用预训练模型,提升跨疾病人群的预测能力

General Demographic Foundation Models for Enhancing Predictive Performance Across Diseases and Populations

  • 构建基于年龄性别表格式数据的通用预训练模型,优化嵌入表示
  • 在多种疾病和人群上验证,显著提升模型判别力与校准性
  • 尤其对依赖年龄性别风险分层的疾病效果明显,适合医疗预测场景

人口统计学特征普遍存在于电子健康记录中,是跨疾病、跨人群最广泛的信息,对临床风险分层和治疗决策至关重要。然而,在模型设计中常被当作辅助信息,其表征学习未受足够重视。本研究探索了针对人口统计学特征(聚焦年龄与性别)的通用人口统计学预训练模型(GDP)的构建。模型在来自不同地理区域、涵盖多种疾病和人群构成的数据集上进行预训练与评估。通过分析排序方法与编码方式的组合,探索了将表格化人口统计输入转化为有效潜在嵌入的架构设计。结果表明,GDP具备跨任务、跨疾病、跨人群的泛化能力。其中,顺序排序显著提升了模型在判别性、校准性及每个决策树分裂点的信息增益表现,尤其在年龄与性别对风险分层贡献较大的疾病中效果突出。即使在人口统计学特征预测价值较低的数据集中,GDP仍能增强其表征重要性,提升下游梯度提升模型中的影响权重。研究显示,针对表格化人口统计学特征的基础模型为提升医疗预测性能提供了有前景的方向。

原文摘要 · Abstract (English)

Demographic attributes are universally present in electronic health records. They are the most widespread information across populations and diseases, and serve as vital predictors in clinical risk stratification and treatment decisions. Despite their significance, these attributes are often treated as auxiliaries in model design, with limited attention being paid to learning their representations. This study explored the development of a General Demographic Pre-trained (GDP) model as a foundational model tailored to demographic attributes, focusing on age and gender. The model is pre-trained and evaluated using datasets with diverse diseases and populations compositions from different geographic regions. The composition of GDP architecture was explored through examining combinations of ordering approaches and encoding methods to transform tabular demographic inputs into effective latent embeddings. Results demonstrate the feasibility of GDP to generalize across task, diseases, and populations. In detailed composition, the sequential ordering substantially improves model performance in discrimination, calibration, and the corresponding information gain at each decision tree split, particularly in diseases where age and gender contribute significantly to risk stratification. Even in datasets where demographic attributes hold relatively low predictive value, GDP enhances the representational importance, increasing their influence in downstream gradient boosting models. The findings suggest that foundation models for tabular demographic attributes offer a promising direction for improving predictive performance in healthcare applications.

医疗预测预训练模型表格式数据人口统计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。