用大模型自动分析死亡预防报告,效率提升百倍且结果可靠。
Automating Thematic Review of Prevention of Future Deaths Reports: Replicating the ONS Child Suicide Study using Large Language Models
- 构建开源文本转表格管道,全自动筛选儿童自杀案例
- 识别出72例儿童自杀报告,是官方数量的近两倍
- 仅8分钟完成全量分析,适合公共卫生政策研究者
英格兰和威尔士的死亡预防报告(PFD)揭示可能引发更多死亡的系统性风险。2025年,英国国家统计局(ONS)对2015年1月至2023年11月期间的儿童自杀类PFD报告进行了全国性主题分析,共识别出37例,全部依赖人工审阅与编码。本文评估了一个完全自动化、开源的“文本到表格”语言模型流程(PFD Toolkit)是否能复现该分析,并检验其效率与可靠性。对2013年7月至2023年11月间发布的4,249份PFD报告进行处理,模型自动筛选出18岁以下个体因自杀致死的报告,并按23个子主题及接收方类别进行编码,复刻了ONS的分类框架。结果共识别出72例儿童自杀报告,接近官方数量的两倍。三位盲法临床医生对144份样本进行独立评审,以共识标注为基准,模型筛查达到显著至近乎完美的一致性(Cohen's κ = 0.82,95% CI: 0.66–0.98,原始一致率91%)。整个端到端脚本运行时间仅8分16秒,将原需数月的工作压缩至分钟级。证明大模型可高效、可靠地复制人工主题审查,实现公共卫生与安全领域的大规模、可重复、及时洞察。PFD Toolkit已开源,供后续研究使用。
原文摘要 · Abstract (English)
Prevention of Future Deaths (PFD) reports, issued by coroners in England and Wales, flag systemic hazards that may lead to further loss of life. Analysis of these reports has previously been constrained by the manual effort required to identify and code relevant cases. In 2025, the Office for National Statistics (ONS) published a national thematic review of child-suicide PFD reports ($\leq$ 18 years), identifying 37 cases from January 2015 to November 2023 - a process based entirely on manual curation and coding. We evaluated whether a fully automated, open source "text-to-table" language-model pipeline (PFD Toolkit) could reproduce the ONS's identification and thematic analysis of child-suicide PFD reports, and assessed gains in efficiency and reliability. All 4,249 PFD reports published from July 2013 to November 2023 were processed via PFD Toolkit's large language model pipelines. Automated screening identified cases where the coroner attributed death to suicide in individuals aged 18 or younger, and eligible reports were coded for recipient category and 23 concern sub-themes, replicating the ONS coding frame. PFD Toolkit identified 72 child-suicide PFD reports - almost twice the ONS count. Three blinded clinicians adjudicated a stratified sample of 144 reports to validate the child-suicide screening. Against the post-consensus clinical annotations, the LLM-based workflow showed substantial to almost-perfect agreement (Cohen's $κ$ = 0.82, 95% CI: 0.66-0.98, raw agreement = 91%). The end-to-end script runtime was 8m 16s, transforming a process that previously took months into one that can be completed in minutes. This demonstrates that automated LLM analysis can reliably and efficiently replicate manual thematic reviews of coronial data, enabling scalable, reproducible, and timely insights for public health and safety. The PFD Toolkit is openly available for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。