用图像识别儿童营养不良,准确高效且适合资源匮乏地区使用。
NutriScreener: Retrieval-Augmented Multi-Pose Graph Attention Network for Malnourishment Screening
- 结合视觉嵌入与知识检索,构建多姿态图注意力模型。
- 在2141个孩子数据上实现0.79召回率、0.82AUC和更低的测量误差。
- 跨区域数据表现提升显著,适合低资源环境部署。
儿童营养不良仍是全球性危机,现有筛查方法耗时且难以推广,阻碍早期干预。本文提出NutriScreener,一种基于CLIP视觉嵌入、类别增强知识检索与上下文感知的检索增强多姿态图注意力网络,可从儿童图像中同时实现营养不良检测与人体测量预测,有效解决泛化性和类别不平衡问题。在临床研究中,医生对其准确率评分4.3/5,效率评分4.6/5,证实其具备实际部署潜力。模型在AnthroVision数据集上训练并测试,额外在跨大陆人群(ARAN及自建CampusPose数据集)上评估,取得0.79召回率、0.82 AUC,并显著降低人体测量均方根误差(RMSE)。跨数据集实验显示,使用人口统计匹配的知识库可使召回率最高提升25%,RMSE降低最多达3.5厘米。该方法为低资源环境下的早期营养不良筛查提供了可扩展且精准的解决方案。
原文摘要 · Abstract (English)
Child malnutrition remains a global crisis, yet existing screening methods are laborious and poorly scalable, hindering early intervention. In this work, we present NutriScreener, a retrieval-augmented, multi-pose graph attention network that combines CLIP-based visual embeddings, class-boosted knowledge retrieval, and context awareness to enable robust malnutrition detection and anthropometric prediction from children's images, simultaneously addressing generalizability and class imbalance. In a clinical study, doctors rated it 4.3/5 for accuracy and 4.6/5 for efficiency, confirming its deployment readiness in low-resource settings. Trained and tested on 2,141 children from AnthroVision and additionally evaluated on diverse cross-continent populations, including ARAN and an in-house collected CampusPose dataset, it achieves 0.79 recall, 0.82 AUC, and significantly lower anthropometric RMSEs, demonstrating reliable measurement in unconstrained pediatric settings. Cross-dataset results show up to 25% recall gain and up to 3.5 cm RMSE reduction using demographically matched knowledge bases. NutriScreener offers a scalable and accurate solution for early malnutrition detection in low-resource environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。