研究大模型在医疗问答中如何被误导,发现断言比虚构证据更难察觉且危害更大。
Untangling the Mechanisms of Misleading Context in Medical Question Answering

- 通过注入虚假证据和断言测试模型对误导信息的敏感性。
- 断言类误导使模型采纳错误答案的频率高出10至27个百分点。
- 只有开放推理轨迹才能有效检测98%的误导决策,而回答本身难以捕捉。
大型语言模型在医疗问答中已达到专家水平,但其输入的上下文可能具有误导性,进而扭曲模型的医学判断。为理解误导上下文如何影响判断,我们考察了模型对上下文的敏感性、误导信息的可披露性、被误导的推理机制以及决策的可监控性。在包含8,627个问题的临床医生评审基准MedMisBench的医疗推理子集上,我们注入两类误导性上下文线索:虚构证据与纯断言。测试了三种推理模型,其中两个暴露完整推理轨迹,一个前沿模型仅输出最终回答。所有模型对断言的敏感度高于虚构证据,采纳断言答案的频率高出10至27分。误导线索在推理轨迹中披露率为81%至98%,但在回答中仅为7%至90%,且断言类线索披露率低于证据类。对未披露轨迹进行重采样显示,虚构证据早期进入并逐步累积影响,而断言则在推理末期直接扭转结论。使用指导提示分析开放推理轨迹时,基于LLM的监控器可在5%假阳性下识别78%的错误决策,远超仅依赖回答的32%。最易导致模型误判的误导信息反而最不被披露,且仅能通过开放推理轨迹可靠检测,而这一能力被主流模型提供商所隐藏。
原文摘要 · Abstract (English)
Large language models now answer medical questions with expert-level performance. However, the context these systems act on can be misleading, and misleading context can corrupt a model's medical judgment. To understand how misleading context corrupts this judgment, we examine the model's susceptibility to the context, disclosure of it, mechanism of corrupted reasoning, and monitorability of the decision. On the medical reasoning subset of MedMisBench, a clinician-reviewed question-answering benchmark of 8,627 questions, we inject two types of misleading context cues, fabricated evidence and a bare assertion. We test three reasoning models, two that expose their full reasoning trace and one frontier model that exposes only its response. All three are more susceptible to the assertion than to the fabricated evidence, adopting the asserted answer 10 to 27 points more often. The misleading cues are disclosed in 81 to 98% of traces but only 7 to 90% of responses, and the assertion is disclosed less often than evidence based cues. Resampling from reasoning traces without disclosure shows the two cues corrupt reasoning differently, evidence entering early and accumulating while the assertion redirects the conclusion near its end. An LLM monitor catches 78% of corrupted decisions at 5% false positives when reading an open model's trace with guidance, against at most 32% from any response. The misleading context that models are most susceptible to is disclosed least, and was caught reliably only from an open reasoning trace, which frontier providers withhold.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。