arXiv:2510.12255cs.CLcs.AI2025-10中稿 · NeurIPS被引 6

测试医疗大模型在多轮对话中的抗干扰能力,发现其易受隐性误导影响。

Shallow Robustness, Deep Vulnerabilities: Multi-Turn Evaluation of Medical LLMs

  • 设计多轮评估框架,区分表面与深层抗干扰能力
  • 多轮交互下模型准确率从91.2%降至13.5%,严重下降
  • 隐性语境误导比直接提示更具破坏性,临床应用需警惕

大语言模型正快速进入医疗临床场景,但其在真实多轮交互下的可靠性仍不清晰。现有评估多基于单轮理想问答,忽视了医疗咨询中常见的冲突输入、误导性上下文和权威影响等复杂情况。本文提出MedQA-Followup框架,系统评估医疗问答中的多轮鲁棒性。方法区分表面鲁棒性(抵抗初始误导)与深层鲁棒性(面对多轮质疑仍保持准确),并引入间接-直接轴,分离情境引导(间接)与明确暗示(直接)。在MedQA数据集上对五种先进LLM进行受控干预测试发现,尽管模型在表面扰动下表现尚可,但在多轮场景中严重失准,如Claude Sonnet 4的准确率从91.2%降至13.5%。反直觉的是,间接情境干预往往比直接建议造成更大准确率下降,暴露出临床部署的重大风险。进一步分析显示模型间差异显著,部分模型在重复干预下持续退化,而另一些则部分恢复甚至提升。结果凸显多轮鲁棒性是医疗LLM安全可靠部署的关键但被忽视维度。

原文摘要 · Abstract (English)

Large language models (LLMs) are rapidly transitioning into medical clinical use, yet their reliability under realistic, multi-turn interactions remains poorly understood. Existing evaluation frameworks typically assess single-turn question answering under idealized conditions, overlooking the complexities of medical consultations where conflicting input, misleading context, and authority influence are common. We introduce MedQA-Followup, a framework for systematically evaluating multi-turn robustness in medical question answering. Our approach distinguishes between shallow robustness (resisting misleading initial context) and deep robustness (maintaining accuracy when answers are challenged across turns), while also introducing an indirect-direct axis that separates contextual framing (indirect) from explicit suggestion (direct). Using controlled interventions on the MedQA dataset, we evaluate five state-of-the-art LLMs and find that while models perform reasonably well under shallow perturbations, they exhibit severe vulnerabilities in multi-turn settings, with accuracy dropping from 91.2% to as low as 13.5% for Claude Sonnet 4. Counterintuitively, indirect, context-based interventions are often more harmful than direct suggestions, yielding larger accuracy drops across models and exposing a significant vulnerability for clinical deployment. Further compounding analyses reveal model differences, with some showing additional performance drops under repeated interventions while others partially recovering or even improving. These findings highlight multi-turn robustness as a critical but underexplored dimension for safe and reliable deployment of medical LLMs.

医疗LLM多轮对话鲁棒性评估模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。