评估大模型在资源匮乏地区应对历史疫情的问答能力,发现其既有潜力也有风险。
Evaluating Large Language Models on Historical Health Crisis Knowledge in Resource-Limited Settings: A Hybrid Multi-Metric Study
- 构建权威来源的问答数据集,多维度评估大模型表现
- 多模型对比显示在疫情知识上存在显著差异和错误
- 适合关注AI辅助公共卫生决策的研究者与政策制定者
大型语言模型(LLMs)在提供健康信息方面具有巨大潜力,但在低资源环境下的可靠性仍不确定。本研究评估了GPT-4、Gemini Pro、Llama~3和Mistral-7B在孟加拉国这一低资源背景下对新冠、登革热、尼帕病毒和基孔肯雅热等健康危机相关问题的回答表现。研究基于权威资料构建了问答数据集,并通过语义相似性、专家-模型交叉评估及自然语言推理(NLI)三种方法进行评估。结果揭示了大模型在呈现流行病学历史与疫情知识方面的优势与局限,凸显其在资源受限环境中支持政策制定的潜力与风险。
原文摘要 · Abstract (English)
Large Language Models (LLMs) offer significant potential for delivering health information. However, their reliability in low-resource contexts remains uncertain. This study evaluates GPT-4, Gemini Pro, Llama~3, and Mistral-7B on health crisis-related enquiries concerning COVID-19, dengue, the Nipah virus, and Chikungunya in the low-resource context of Bangladesh. We constructed a question--answer dataset from authoritative sources and assessed model outputs through semantic similarity, expert-model cross-evaluation, and Natural Language Inference (NLI). Findings highlight both the strengths and limitations of LLMs in representing epidemiological history and health crisis knowledge, underscoring their promise and risks for informing policy in resource-constrained environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。