arXiv:2601.20987cs.LG2026-01

预训练模型让发展迟缓检测在数据极少的国家也能用。

Pre-trained Encoders for Global Child Development: Transfer Learning Enables Deployment in Data-Scarce Settings

  • 用全球44国35万+儿童数据预训练编码器,支持少样本迁移。
  • 仅需50个样本,模型AUC达0.65,比传统方法高8-12%。
  • 零样本部署到新国家,最高AUC达0.84,适合资源匮乏地区。

每年有大量儿童遭遇可预防的发展迟缓,但机器学习在新国家的部署受限于数据瓶颈:可靠模型需数千样本,而新项目初始数据常不足100。本文提出首个面向全球儿童发展的预训练编码器,基于联合国儿童基金会(UNICEF)调查数据,在44个国家的357,709名儿童上训练。仅用50个样本,该编码器平均AUC达0.65(95%置信区间:0.56–0.72),在各区域均比冷启动梯度提升模型的0.61高出8%-12%。当样本量增至500时,AUC达到0.73。在未见过的国家实现零样本部署,最高AUC达0.84。我们引入迁移学习理论边界,解释预训练多样性如何促进少样本泛化。结果表明,预训练编码器可显著提升资源匮乏环境中对可持续发展目标4.2.1的监测可行性。

原文摘要 · Abstract (English)

A large number of children experience preventable developmental delays each year, yet the deployment of machine learning in new countries has been stymied by a data bottleneck: reliable models require thousands of samples, while new programs begin with fewer than 100. We introduce the first pre-trained encoder for global child development, trained on 357,709 children across 44 countries using UNICEF survey data. With only 50 training samples, the pre-trained encoder achieves an average AUC of 0.65 (95% CI: 0.56-0.72), outperforming cold-start gradient boosting at 0.61 by 8-12% across regions. At N=500, the encoder achieves an AUC of 0.73. Zero-shot deployment to unseen countries achieves AUCs up to 0.84. We apply a transfer learning bound to explain why pre-training diversity enables few-shot generalization. These results establish that pre-trained encoders can transform the feasibility of ML for SDG 4.2.1 monitoring in resource-constrained settings.

儿童发展迁移学习少样本预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。