检验医学推理链是否真实反映模型思考,发现多数链与答案脱节。
Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought
- 设计临床相关扰动操作,测试推理链与答案的关联性。
- 72.9%的扰动下链不变化而答案翻转,说明链不可信。
- 适用于评估医疗大模型推理真实性,尤其适合关注可解释性的研究者。
临床医生将链式思维(CoT)视为医疗推理的证据,但其实际作用极少被验证。通用领域CoT忠实性检测忽略临床代价,而医学大模型评估则将链视为黑箱。本文提出医疗扰动审计:设计30种临床动机的编辑操作(如严重程度反转、否定翻转、人口统计替换、证据移除),结合链更新与答案翻转联合分析,对14个LLM在四个医学QA基准上进行测试。三重独立检验一致显示:在临床有意义的破坏性编辑下,链解耦率(CDR)达72.9%;链被篡改后准确率不变;去除CoT提示也不降低准确率。两名认证医师重新标注197个扰动问题,98.5%仍保有合理金标准。该模式跨医学微调、推理微调及模型规模均成立;对闭源模型,仅从答案侧信号亦观察到相同解耦现象。本框架与CDR为评估医学CoT是否忠实提供可复用标尺。
原文摘要 · Abstract (English)
Clinicians read chain-of-thought (CoT) rationales as evidence of medical reasoning, but whether the visible chain plays that role is rarely tested. General-domain CoT-faithfulness probes ignore clinical cost, and medical LLM evaluations treat the chain as a black box. We close this gap with a medical perturbation audit: a 30-operator battery edits both the chain and the question with clinically motivated operators (severity reversal, negation flip, demographic swap, evidence ablation), paired with a chain-update times answer-flip joint analysis that classifies each model by its failure mode. Applied to 14 LLMs on four medical QA benchmarks, three independent tests converge: the Chain-Decoupling Rate (CDR; chain does not register the edit and the answer does not flip) is 72.9% panel-wide on clinically meaningful destructive edits, chain corruption leaves accuracy unchanged, and removing CoT prompting does not reduce accuracy. Two board-certified clinicians re-annotate N=197 perturbed questions; 98.5% leave the gold defensible. The pattern holds across medical and reasoning fine-tuning and scale; on the closed-source tier, where the chain text is unavailable, the answer-side signals are consistent with the same decoupling. Our framework and CDR provide a reusable yardstick for auditing whether medical CoT is faithful or merely documentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。