用上下文信息提升大模型对热带病的诊断能力,让AI更贴近真实临床场景。
Contextual Evaluation of Large Language Models for Classifying Tropical and Infectious Diseases
- 引入性别、位置、风险因素等上下文信息增强医学提示数据
- 在1.1万+个提示上测试,发现上下文使大模型表现接近人类专家
- 开发了TRINDs-LM原型工具,供研究者探索上下文如何影响AI输出
尽管大语言模型在医疗问答中展现出潜力,但针对热带和传染病领域的研究仍有限。我们基于开源的热带与传染病(TRINDs)数据集,扩展其内容,加入人口统计学与语义上的临床及消费者增强信息,形成超过11000个提示样本。我们在这些数据上评估通用型与医学专用大模型的表现,并与人类专家进行对比。通过系统性实验,我们证明了包括性别、地理位置、风险因素在内的上下文信息对于获得最优大模型响应至关重要。最后,我们构建了TRINDs-LM原型工具,为研究者提供一个交互平台,以探索上下文如何影响大模型在健康领域的输出。
原文摘要 · Abstract (English)
While large language models (LLMs) have shown promise for medical question answering, there is limited work focused on tropical and infectious disease-specific exploration. We build on an opensource tropical and infectious diseases (TRINDs) dataset, expanding it to include demographic and semantic clinical and consumer augmentations yielding 11000+ prompts. We evaluate LLM performance on these, comparing generalist and medical LLMs, as well as LLM outcomes to human experts. We demonstrate through systematic experimentation, the benefit of contextual information such as demographics, location, gender, risk factors for optimal LLM response. Finally we develop a prototype of TRINDs-LM, a research tool that provides a playground to navigate how context impacts LLM outputs for health.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。