评测大模型对癌症患者含错误假设问题的回应能力,发现其纠错率不足43%。
Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions
- 构建专家验证的对抗性数据集Cancer-Myth,包含585个带错误预设的问题
- 主流大模型纠错率均不超过43%,存在严重医疗风险
- 提示词优化可提升至80%准确率,但误判率高达41%且影响其他任务
癌症患者越来越多地使用大语言模型获取医疗信息,因此评估模型处理复杂个性化问题的能力至关重要。然而现有医疗基准多聚焦于医学考试或消费者搜索问题,未涵盖真实患者带有个人细节的问题。本文由三位血液肿瘤学医生评估来自真实患者的癌症相关问题。尽管模型回答总体准确,但常无法识别或回应问题中的错误预设,危及安全医疗决策。为此,我们构建了专家验证的对抗性数据集Cancer-Myth,含585个带虚假预设的癌症问题。在该基准上,前沿大模型(包括GPT-5、Gemini-2.5-Pro、Claude-4-Sonnet)纠正错误预设的比例均不超过43%。为进一步研究缓解策略,我们建立150个无虚假预设的Cancer-Myth-NFP子集。结果表明,典型提示优化方法(如结合GEPA优化的警示提示)可将Cancer-Myth准确率提升至80%,但导致41%的NFP问题被误判,并在其他医疗基准上造成10%相对性能下降。研究揭示了大模型在可靠性上的关键缺陷,证明仅靠提示无法可靠解决虚假预设问题,亟需更稳健的医疗AI防护机制。
原文摘要 · Abstract (English)
Cancer patients are increasingly turning to large language models (LLMs) for medical information, making it critical to assess how well these models handle complex, personalized questions. However, current medical benchmarks focus on medical exams or consumer-searched questions and do not evaluate LLMs on real patient questions with patient details. In this paper, we first have three hematology-oncology physicians evaluate cancer-related questions drawn from real patients. While LLM responses are generally accurate, the models frequently fail to recognize or address false presuppositions in the questions, posing risks to safe medical decision-making. To study this limitation systematically, we introduce Cancer-Myth, an expert-verified adversarial dataset of 585 cancer-related questions with false presuppositions. On this benchmark, no frontier LLM -- including GPT-5, Gemini-2.5-Pro, and Claude-4-Sonnet -- corrects these false presuppositions more than $43\%$ of the time. To study mitigation strategies, we further construct a 150-question Cancer-Myth-NFP set, in which physicians confirm the absence of false presuppositions. We find typical mitigation strategies, such as adding precautionary prompts with GEPA optimization, can raise accuracy on Cancer-Myth to $80\%$, but at the cost of misidentifying presuppositions in $41\%$ of Cancer-Myth-NFP questions and causing a $10\%$ relative performance drop on other medical benchmarks. These findings highlight a critical gap in the reliability of LLMs, show that prompting alone is not a reliable remedy for false presuppositions, and underscore the need for more robust safeguards in medical AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。