测试大模型能否科学分配儿童铅暴露检测资源,结果发现其表现不佳。
Can LLMs Help Allocate Public Health Resources? A Case Study on Childhood Lead Testing
- 用多指标综合评分法生成优先级,指导三城136个社区的检测资源分配
- 大模型平均准确率仅0.46,最高0.66,常误判高风险区域
- 适合关注医疗资源分配与AI可靠性交叉研究的读者
公共卫生机构在资源有限的情况下,难以识别儿童铅暴露的高风险社区。为此,我们构建了一个综合未检测儿童比例、高血铅流行率及公共卫生覆盖模式的优先级评分体系,用于优化芝加哥、纽约市和华盛顿特区共136个社区的资源分配。利用需整合多重脆弱性指标并基于实证证据决策的任务,评估具备代理推理与深度研究能力的大语言模型(LLMs)在结构化分配场景下的表现。任务要求各城市将1000个检测套装按社区脆弱性指标分配。结果显示,大模型频繁忽略铅暴露率最高、未检测儿童占比最大的区域(如芝加哥西恩格尔伍德),反而向低优先级区域(如纽约市亨茨点)分配过多资源。整体准确率平均为0.46,最高达0.66(使用ChatGPT 5 Deep Research)。尽管宣传具备深度研究能力,大模型仍暴露出信息检索与基于证据推理的根本缺陷,常引用过时数据,让非实证叙述取代量化指标。
原文摘要 · Abstract (English)
Public health agencies face critical challenges in identifying high-risk neighborhoods for childhood lead exposure with limited resources for outreach and intervention programs. To address this, we develop a Priority Score integrating untested children proportions, elevated blood lead prevalence, and public health coverage patterns to support optimized resource allocation decisions across 136 neighborhoods in Chicago, New York City, and Washington, D.C. We leverage these allocation tasks, which require integrating multiple vulnerability indicators and interpreting empirical evidence, to evaluate whether large language models (LLMs) with agentic reasoning and deep research capabilities can effectively allocate public health resources when presented with structured allocation scenarios. LLMs were tasked with distributing 1,000 test kits within each city based on neighborhood vulnerability indicators. Results reveal significant limitations: LLMs frequently overlooked neighborhoods with highest lead prevalence and largest proportions of untested children, such as West Englewood in Chicago, while allocating disproportionate resources to lower-priority areas like Hunts Point in New York City. Overall accuracy averaged 0.46, reaching a maximum of 0.66 with ChatGPT 5 Deep Research. Despite their marketed deep research capabilities, LLMs struggled with fundamental limitations in information retrieval and evidence-based reasoning, frequently citing outdated data and allowing non-empirical narratives about neighborhood conditions to override quantitative vulnerability indicators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。