arXiv:2607.29082cs.CL2026-07中稿 · AIME 2026 Workshop

用大模型零样本预测儿童发育迟缓,效果不输传统方法,但城乡和贫富差距影响公平性。

Can Zero-Shot LLMs Predict Child Malnutrition? A Fairness and Temporal Robustness Study

  • 将问卷数据转为提示词输入GPT-4o-mini,实现零样本预测
  • 准确率接近监督模型,对发育迟缓病例识别率更高
  • 适合关注医疗公平性与长期预测稳定性的研究者

儿童营养不良仍是低收入和中等收入国家的重大公共卫生挑战,尤其在南亚地区,早期识别高危儿童对及时干预和资源分配至关重要。本研究评估了预训练大语言模型(LLM)在零样本设置下,利用人口健康调查数据预测儿童发育迟缓的可行性、公平性与时间鲁棒性。基于2007至2022年间孟加拉国人口健康调查(BDHS)数据,将母亲、儿童、医疗及家庭特征转化为语义可解释的提示词表示,评估GPT-4o-mini在零样本下的发育迟缓预测表现,并与随机森林基线对比,同时分析不同人口学和社会经济群体间的公平性差异以及跨调查波次的时间鲁棒性。结果表明,使用GPT-4o-mini进行零样本推理的平衡准确率与监督基线相当,对发育迟缓病例的敏感性显著更高,在儿童性别组间表现相对一致,且在不同年份的调查数据中保持稳定的预测行为;然而,在居住地与家庭财富类别之间仍存在重要公平性差异,提示在公共卫生预测中部署基础模型前需进一步深入研究。

原文摘要 · Abstract (English)

Child malnutrition remains a major public health challenge in low- and middle-income countries, particularly in South Asia, where early identification of vulnerable children is critical for timely intervention and resource allocation. This study aims to evaluate the feasibility, fairness, and temporal robustness of using a pretrained large language model (LLM) in a zero-shot setting for child stunting prediction using population health survey data. Using Bangladesh Demographic and Health Survey (BDHS) data collected between 2007 and 2022, we transformed maternal, child, healthcare, and household characteristics into semantically interpretable prompt-based representations and evaluated GPT-4o-mini for zero-shot stunting prediction, comparing its performance against a random forest baseline and assessing fairness across demographic and socioeconomic groups as well as temporal robustness across survey waves. The results demonstrate that zero-shot inference using GPT-4o-mini achieved comparable balanced accuracy to the supervised baseline while exhibiting substantially higher sensitivity for identifying stunting cases, relatively consistent performance across child sex groups, and stable predictive behaviour across BDHS waves; however, important fairness disparities were observed across residence and household wealth categories, highlighting the need for further investigation before deployment of foundation models in public health prediction settings.

大模型应用医疗公平零样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。