检测网页医疗大模型的幻觉与滥用,发现近三成模型存在事实错误。
Do No Harm? Hallucination and Actor-Level Abuse in Web-Deployed Medical Large Language Models
- 构建双框架评估医疗模型幻觉与政策违规,涵盖1500个模型
- 25%-30%模型事实准确率低,超一半带功能模型隐私披露不足
- 适合关注医疗AI安全的研究者与平台开发者
医疗大语言模型(包括定制化MedGPT和开源模型)正被广泛部署于网络平台提供临床建议。然而,这些模型存在幻觉、政策不合规及设计不安全等风险。我们对6,233个MedGPT进行了大规模评估,选取1,500个分层样本,并对比10个开源模型。提出两个评估框架:MedGPT-HEval用于幻觉检测,以及基于LLM的策略违规与开发者意图分析管道。结果显示,25%-30%的MedGPT存在低事实准确性,底层与中端模型风险最高;33.6%-54.3%违反操作阈值,57.06%的可执行动作模型缺乏充分隐私披露。尽管相比开源模型,MedGPT在事实准确性和语义对齐上表现更优,但开源模型更稳定。研究揭示了幻觉与合规性方面的系统性缺口,强调需多指标评估与更强防护机制。我们发布了HAA-MedGPT,一个结构化数据集,支持未来医疗大模型安全研究。
原文摘要 · Abstract (English)
Medical large language models (LLMs), including custom medical GPTs (MedGPTs) and open-source models, are increasingly deployed on web platforms to provide clinical guidance. However, they pose risks of hallucination, policy noncompliance, and unsafe design. We conduct a large-scale assessment of 6,233 MedGPTs, evaluating a stratified sample of 1,500, together with 10 open-source LLMs. We introduce two frameworks: MedGPT-HEval for hallucination detection and an LLM-based pipeline for assessing policy violations and developer intent. Our results show that 25-30% of MedGPTs exhibit low factual accuracy, with bottom- and middle-tier models at highest risk; 33.6-54.3% violate operational thresholds, and 57.06% of Action-enabled models lack adequate privacy disclosures. Compared with open-source models, MedGPTs achieve higher factual accuracy and semantic alignment, though open-source models are more stable. These results reveal systemic gaps in hallucination and compliance, highlighting the need for multi-metric evaluation and stronger safeguards. We release HAA-MedGPT, a structured dataset that supports future research on the safety of web-facing medical LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。