用少量数据预测儿童贫血,基础模型比传统方法更准。
Few-shot Cross-country Generalization of Tabular Machine Learning and Foundation Models for Childhood Anemia Prediction under Distribution Shift

- 用Transformer结构的表格基础模型,应对数据稀缺场景
- 在少样本下表现更好,AUC最高达0.76,误差更低
- 适合资源匮乏地区医疗预测,尤其关注低数据环境
全球约40%的6-59月龄儿童患有贫血,成因多样导致模型泛化困难。本研究在非洲、亚洲、拉丁美洲、高加索及中东16个国家的DHS数据(n=68,856)上,对比了逻辑回归、XGBoost、LightGBM与TabPFN v2.6在跨国家、低数据条件下的表现。采用AUC-ROC、Brier得分和ECE评估性能,通过留一国剔除(LOCO)、反向留一国(reverse-LOCO)及少样本设置检验泛化能力。子组分析涵盖性别、年龄、居住地、母亲教育和财富水平。特征重要性使用SHAP估算。结果显示,在样本<200的低数据环境下,TabPFN优于传统模型,判别能力更强且校准更优;跨国家平均Brier得分为0.042,ECE为0.203。全数据条件下AUC-ROC为0.59-0.76,模型间差异≤0.05。LOCO表现稳定(0.58-0.69),受国家背景影响;reverse-LOCO显示迁移不对称。各子组表现一致,无系统性人口偏差。SHAP识别出儿童年龄、海拔、身高/年龄z分数为主要预测因子,其次为财富与母亲教育。贫血预测效果更多取决于人群差异而非模型选择。在低资源场景中,TabPFN凭借更强判别力与校准性,展现出作为基础模型在公共卫生预测中的潜力。
原文摘要 · Abstract (English)
Childhood anemia affects around 40% of children aged 6-59 months globally and arises from heterogeneous factors, limiting model generalizability. We evaluate a transformer-based tabular foundation model against classical supervised methods under cross-country and data-scarce settings. We used DHS data from 16 countries across Africa, Asia, Latin America, the Caucasus, and the Middle East (n=68,856). We compared Logistic Regression, XGBoost, LightGBM, and TabPFN v2.6. Performance was assessed using AUC-ROC, Brier score, and ECE. Generalization was evaluated using leave-one-country-out (LOCO), reverse-LOCO, and few-shot settings. Subgroup analyses included sex, age, residence, maternal education, and wealth. Feature importance was estimated using SHAP. TabPFN outperformed classical models in low-data regimes (<200 samples), showing higher discrimination and better calibration. Across countries, it achieved the lowest Brier score (0.042) and ECE (0.203). Under full-data settings, AUC-ROC ranged from 0.59-0.76 with small between-model differences ($\leq 0.05$). LOCO performance was stable (0.58-0.69), driven by country context. Reverse-LOCO showed asymmetric transferability. Subgroup performance was consistent with no systematic demographic bias. SHAP identified child age, altitude, and height-for-age z-score as dominant predictors, followed by wealth and maternal education. Performance in childhood anemia prediction is driven more by population variation than model choice. TabPFN provides advantages in low-resource settings through improved discrimination and calibration, highlighting foundation models as promising tools for data-scarce global health prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。