测试大模型在错误医学信息干扰下的判断稳定性,发现其准确率大幅下降。
Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

- 构建新评测集MedMisBench,注入误导性上下文测试模型鲁棒性
- 平均准确率从71.1%降至38.0%,攻击成功率高达51.5%
- 权威型虚假信息和例外污染类误导最致命,临床专家认出38.2%有风险
大型语言模型(LLMs)在医学执照考试中已达到专家水平,促使人们认为高分即代表安全可靠的医疗判断。然而我们发现这一假设脆弱:当在原本正确回答的问题中注入误导性上下文时,模型会放弃正确答案。我们将这种在对抗性上下文中保持正确判断的能力称为认知韧性,并提出MedMisBench来衡量它。该评测集包含10,932个医学问题项和48,889对误导性上下文-选项组合,涵盖医学推理、代理能力与患者旅程评估。在11种模型配置下,平均准确率从原始问题的71.1%下降至聚焦误导上下文下的38.0%,攻击成功率达51.5%。最具破坏性的误导形式为形式化、规则化的虚构内容:以权威框架呈现的谎言攻击成功率达69.5%,例外污染类陈述达64.1%。由7个国家14名临床专家组成的评审团在38.2%的案例中识别出严重潜在危害。MedMisBench揭示了当前医学领域大模型评估的结构性盲区:现有基准仅衡量模型的知识掌握程度,而未考察其在误导情境下维持正确判断的能力。
原文摘要 · Abstract (English)
Large language models (LLMs) now reach expert-level scores on medical licensing exams, encouraging the assumption that high scores imply safe medical judgment while patients increasingly use them for health advice. We show this assumption is fragile: when misleading context is injected into questions that LLMs originally answer correctly, they abandon the correct answer. We call the ability to maintain correct judgment under adversarial context epistemic resilience, and introduce MedMisBench to measure it. MedMisBench contains 10,932 medical question items and 48,889 misleading context-option pairs spanning medical reasoning, agentic capability, and patient-journey evaluation. Across 11 model configurations, mean accuracy falls from 71.1% on original questions to 38.0% under focused misleading context, with 51.5% attack success. The most damaging injections are formal, rule-like fabrications: authority-framed falsehoods reach 69.5% attack success and exception-poisoning claims reach 64.1%. A 14-member clinical panel from 7 countries identified serious potential harm in 38.2% of reviewed cases. MedMisBench exposes a structural blind spot in LLM evaluation in medical settings: existing benchmarks measure what models know, but not whether they preserve correct medical judgment under misleading context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。