医学推理模型虽答对率提升,但推理步骤错误率反而上升。
Better Accuracies, Worse Reasoning: A Step-Level Audit of Medical Chain-of-Thought Distillation
- 用大模型推理轨迹训练小模型,模仿专家思考过程
- 答案准确率从74.7%升至84.4%,但推理步骤错误率从30.6%升至50.3%
- 当前评估方法无法发现推理质量退化,临床应用需警惕
链式思维(CoT)蒸馏通过让小模型模仿大模型的推理过程来提升性能,但通常仅以最终答案指标(如准确率)评估。在医学问答任务中,基于DeepSeek-V3系列教师训练的Qwen3-8B学生模型,在MedQA-USMLE数据集上答案准确率从74.7%提升至84.4%(SC@64),预期校准误差(ECE)从0.096降至0.034。然而,在采用Kimi-K2.6风格盲测的LLM评审下,非放弃步骤的错误率从30.6%升至50.3%。该现象在不同评估者、教师能力、学生规模与架构、医学基准及风格、分段方式和答案正确性控制下均持续存在。一位临床专家进行的150步盲审验证了相同趋势。边界分析表明,当答案过简导致推理依据不足,且小模型能模仿专家形式但缺乏真实支撑时,风险显现。现有答案指标与整体置信度统计无法揭示这一变化。因此,仅依赖答案级评估会掩盖推理质量下降,释放或复用此类推理轨迹时需谨慎。
原文摘要 · Abstract (English)
Chain-of-thought (CoT) distillation trains a smaller model to imitate a teacher's reasoning trace, but it is typically evaluated by final-answer metrics including accuracy. We ask whether gains in answer quality are accompanied by improvements in the trace. In medical QA, where short answer options can leave a richer clinical justification under-specified, a Qwen3-8B student distilled from a DeepSeek-V3-family teacher improves on MedQA-USMLE answer metrics (SC@64 74.7% to 84.4%; expected calibration error (ECE) 0.096 to 0.034). Yet under a Kimi-K2.6 style-blind LLM-judge audit, its error rate over non-abstained steps rises from 30.6% to 50.3%. In this primary medical setting, answer quality and trace factuality move in opposite directions. This before--after pattern persists across evaluators, teacher strengths, student scales and families, medical benchmarks, and style, segmentation, and answer-correctness controls. A 150-step blinded audit by a clinical expert reproduces the same ordering. Boundary checks narrow the scope of the claim: the risk appears when a compact answer under-constrains the rationale and a capable student can imitate expert-like form without reliably grounding each local claim. Standard answer metrics and aggregate hedging rates do not reveal the shift. When such traces are released or reused, answer-level metrics alone are insufficient.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。