用大模型分析死亡访谈文本,提升病因判断准确率
Leveraging Language Models and Machine Learning in Verbal Autopsy Analysis
- 用预训练语言模型处理未结构化访谈文本进行病因分类
- 仅用访谈文本就超越传统问答算法,尤其擅长识别慢病
- 建议改进问卷设计,重视叙述性信息的价值
在缺乏出生死亡登记的国家,口头尸检(VA)是估算死因和制定政策的关键工具。VA通过向知情人询问死亡前的状况,获取非结构化叙述和结构化问题回答。现有自动化死因分类方法仅使用问题数据,忽略叙述内容。本文利用南非实证数据,研究如何用预训练语言模型(PLMs)和机器学习技术处理VA叙述文本。结果表明,仅使用叙述文本,经任务微调的基于Transformer的PLM在个体与群体层面均优于主流仅用问题的算法,尤其在识别非传染性疾病方面表现突出。我们探索了融合叙述与问题的多模态策略,发现联合使用能进一步提升性能,证明两种模态各有独特价值。同时,我们分析了医生对信息充分性的评估,发现不同年龄和死因下信息充分性存在差异,且模型与医生的分类准确率均受此影响。本研究推动了自然语言处理、流行病学与全球健康交叉领域的进展,证实叙述文本对死因分类具有重要价值。研究呼吁收集更多来自多样化场景的高质量数据,以支持未来模型训练,并为重新设计VA问卷提供依据。
原文摘要 · Abstract (English)
In countries without civil registration and vital statistics, verbal autopsy (VA) is a critical tool for estimating cause of death (COD) and inform policy priorities. In VA, interviewers ask proximal informants for details on the circumstances preceding a death, in the form of unstructured narratives and structured questions. Existing automated VA cause classification algorithms only use the questions and ignore the information in the narratives. In this thesis, we investigate how the VA narrative can be used for automated COD classification using pretrained language models (PLMs) and machine learning (ML) techniques. Using empirical data from South Africa, we demonstrate that with the narrative alone, transformer-based PLMs with task-specific fine-tuning outperform leading question-only algorithms at both the individual and population levels, particularly in identifying non-communicable diseases. We explore various multimodal fusion strategies combining narratives and questions in unified frameworks. Multimodal approaches further improve performance in COD classification, confirming that each modality has unique contributions and may capture valuable information that is not present in the other modality. We also characterize physician-perceived information sufficiency in VA. We describe variations in sufficiency levels by age and COD and demonstrate that classification accuracy is affected by sufficiency for both physicians and models. Overall, this thesis advances the growing body of knowledge at the intersection of natural language processing, epidemiology, and global health. It demonstrates the value of narrative in enhancing COD classification. Our findings underscore the need for more high-quality data from more diverse settings to use in training and fine-tuning PLM/ML methods, and offer valuable insights to guide the rethinking and redesign of the VA instrument and interview.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。