检测大模型在医学教材问答中的幻觉现象,发现70B模型仍有近20%错误回答。
Quantifying Hallucinations in Language Language Models on Medical Textbooks
- 用教材原文作证据源,测试大模型医学问答幻觉率。
- LLaMA-70B-Instruct在测试中幻觉率达19.7%,但多数回答看似合理。
- 医生评价显示幻觉少的模型更实用,专家审核成本高。
幻觉是大语言模型生成与事实不符或无依据内容的现象,目前尚无有效缓解方法。现有医学问答评测很少基于固定证据源进行。本文研究了基于教材的医学问答中幻觉的发生频率及不同模型响应差异。第一项实验评估了主流开源大模型 LLaMA-70B-Instruct 在闭源零样本提示下的幻觉率,结果发现其在提供段落证据的情况下,仍有19.7%(95%置信区间18.6至20.7)的答案存在幻觉,而98.8%的回答被认为具有最高合理性。第二项实验对比了多个模型的幻觉率与临床医生偏好,结果显示幻觉率越低,有用性评分越高(ρ = -0.71, p = 0.058)。医生对实验一和实验二的评价一致性高(加权κ = 0.92;τ_b = 0.06 至 0.18,κ = 0.57 至 0.61)。研究表明,当前所有规模与架构的大模型均不适合未经监督的临床部署,人类专家审查既必要又为主要成本来源。
原文摘要 · Abstract (English)
Hallucinations, the tendency for large language models to provide responses with factually incorrect and unsupported claims, is a serious problem within natural language processing for which we do not yet have an effective solution to mitigate against. Existing benchmarks for medical QA rarely evaluate this behavior against a fixed evidence source. We ask how often hallucinations occur on textbook-grounded QA and how responses to medical QA prompts vary across models. We conduct two experiments, the first experiment to determine the prevalence of hallucinations for a prominent open source large language model (LLaMA-70B-Instruct) in medical QA given closed-source zero-shot prompts, and the second experiment to determine the prevalence of hallucinations and clinician preference to model responses. We observed, in experiment one, with the passages provided, LLaMA-70B-Instruct hallucinated in 19.7\% of answers (95\% CI 18.6 to 20.7) even though 98.8\% of prompt responses received maximal plausibility, and observed in experiment two, across models, lower hallucination rates aligned with higher usefulness scores ($ρ=-0.71$, $p=0.058$). Clinicians produced high agreement (quadratic weighted $κ=0.92$) and ($τ_b=0.06$ to $0.18$, $κ=0.57$ to $0.61$) for experiments 1 and 2 respectively. Our findings indicate that, across all scales and architectures tested, current large language models remain unfit for unsupervised clinical deployment, and that human expert oversight is both necessary and the dominant cost driver.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。