针对胰腺癌临床场景,构建了3130条真实患者问题的评估基准,揭示大模型在准确性与幻觉上的显著差异。
PanCanBench: A Comprehensive Benchmark for Evaluating Large Language Models in Pancreatic Oncology
- 基于真实患者问题与专家评审,构建胰腺癌专用评估基准PanCanBench
- 模型幻觉率最高达53.8%,新推理模型未显著提升事实正确性
- 强调真实临床场景评估必要性,适合医疗AI安全研究者参考
大语言模型在标准化考试中表现接近专家水平,但多项选择题准确率难以反映其真实临床价值与安全性。随着患者和医生越来越多地依赖大模型处理胰腺癌等复杂疾病,评估需超越通用医学知识。现有框架如HealthBench依赖模拟查询,缺乏疾病特异性深度。高评分未必代表事实正确,凸显幻觉检测的必要性。我们构建了人机协同流程,基于胰腺癌行动网络(PanCAN)去标识化患者问题生成专家评分标准。由此产生的基准PanCanBench包含282个真实患者问题,共3130条具体评估标准。我们使用大模型作为裁判框架评估22个专有及开源模型,测量临床完整性、事实准确性和网页搜索整合能力。模型在评分标准下的完整性得分从46.5%至82.3%不等。事实错误普遍存在,幻觉率(响应中至少含一个事实错误的比例)从Gemini-2.5 Pro和GPT-4o的6.0%到Llama-3.1-8B的53.8%。值得注意的是,较新的推理优化模型并未一致提升事实性:虽然o3获得最高评分,但其错误频率高于其他GPT系列模型。网页搜索整合也未必然带来更好结果:Gemini-2.5 Pro启用搜索后平均分从66.8%降至63.9%,GPT-5从73.8%降至72.8%。合成生成的评分标准使平均分虚高17.9分,但相对排名基本保持一致。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved expert-level performance on standardized examinations, yet multiple-choice accuracy poorly reflects real-world clinical utility and safety. As patients and clinicians increasingly use LLMs for guidance on complex conditions such as pancreatic cancer, evaluation must extend beyond general medical knowledge. Existing frameworks, such as HealthBench, rely on simulated queries and lack disease-specific depth. Moreover, high rubric-based scores do not ensure factual correctness, underscoring the need to assess hallucinations. We developed a human-in-the-loop pipeline to create expert rubrics for de-identified patient questions from the Pancreatic Cancer Action Network (PanCAN). The resulting benchmark, PanCanBench, includes 3,130 question-specific criteria across 282 authentic patient questions. We evaluated 22 proprietary and open-source LLMs using an LLM-as-a-judge framework, measuring clinical completeness, factual accuracy, and web-search integration. Models showed substantial variation in rubric-based completeness, with scores ranging from 46.5% to 82.3%. Factual errors were common, with hallucination rates (the percentages of responses containing at least one factual error) ranging from 6.0% for Gemini-2.5 Pro and GPT-4o to 53.8% for Llama-3.1-8B. Importantly, newer reasoning-optimized models did not consistently improve factuality: although o3 achieved the highest rubric score, it produced inaccuracies more frequently than other GPT-family models. Web-search integration did not inherently guarantee better responses. The average score changed from 66.8% to 63.9% for Gemini-2.5 Pro and from 73.8% to 72.8% for GPT-5 when web search was enabled. Synthetic AI-generated rubrics inflated absolute scores by 17.9 points on average while generally maintaining similar relative ranking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。