用大模型辅助尸检,提升资源匮乏地区死因判断准确率。
LAVA: Language Model Assisted Verbal Autopsy for Cause-of-Death Determination
- 结合大模型与传统算法,构建端到端死因预测流程。
- 大模型在成人、儿童、新生儿中的准确率分别达48.6%、50.5%、53.5%。
- 无需复杂训练,可直接应用,适合低资源公共卫生场景。
在缺乏医疗死亡证明的资源匮乏地区,口头尸检(VA)是估算死因的重要工具。本研究提出LA-VA原型流程,融合大语言模型(LLMs)、传统算法与基于嵌入的分类方法,以提升死因预测性能。基于人口健康指标研究联盟(PHMRC)数据集,覆盖成人(7,580例)、儿童(1,960例)和新生儿(2,438例)三类人群,评估了GPT-5预测、LCVA基线、文本嵌入及元学习集成等多种方法。结果表明,GPT-5在各类别中表现最优,平均测试站点准确率分别为48.6%(成人)、50.5%(儿童)和53.5%(新生儿),较传统统计机器学习基线高出5%-10%。研究提示,仅使用现成的大模型即可显著提升口头尸检精度,对全球低资源环境下的疾病监测具有重要意义。
原文摘要 · Abstract (English)
Verbal autopsy (VA) is a critical tool for estimating causes of death in resource-limited settings where medical certification is unavailable. This study presents LA-VA, a proof-of-concept pipeline that combines Large Language Models (LLMs) with traditional algorithmic approaches and embedding-based classification for improved cause-of-death prediction. Using the Population Health Metrics Research Consortium (PHMRC) dataset across three age categories (Adult: 7,580; Child: 1,960; Neonate: 2,438), we evaluate multiple approaches: GPT-5 predictions, LCVA baseline, text embeddings, and meta-learner ensembles. Our results demonstrate that GPT-5 achieves the highest individual performance with average test site accuracies of 48.6% (Adult), 50.5% (Child), and 53.5% (Neonate), outperforming traditional statistical machine learning baselines by 5-10%. Our findings suggest that simple off-the-shelf LLM-assisted approaches could substantially improve verbal autopsy accuracy, with important implications for global health surveillance in low-resource settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。