让AI医生的思考过程可被评估和优化,提升医疗决策可信度。
Constructing and Evaluating Clinical Reasoning Trajectories for Medical Agent

- 构建多步推理链并从五个维度量化其质量
- 引入错误注入验证关键步骤对结果的影响,精度提升0.029~0.041
- 仅少数核心步骤影响整体质量,适合医疗AI研发与临床应用
当前医学AI评估仍以最终答案为中心,忽略中间推理质量。在临床场景中,通过虚构证据或逻辑混乱得出正确答案同样危险。本文提出MedTraj框架,将推理轨迹作为关键对象进行构建、评估与优化。该流程从医学推理源生成结构化多步推理链,解析为临床观察、证据、编号推理步骤与结论,并在连贯性、证据支持、幻觉、完整性与可追溯性五个维度评分。通过控制性错误注入,建立特定推理失败与质量下降之间的因果关系。基于边际贡献分析,识别影响轨迹质量的关键推理步骤。最后,质量加权上下文学习将评估反馈引入模型推理,使其从强弱推理示例中学习。在CareQA、PubMedQA和CECMed上实验表明,轨迹上下文显著提升推理连贯性,相较零样本基线提升0.029至0.041;在CECMed上,正确率接近翻倍,幻觉率降低87%。边际分析显示,少数推理步骤承载主要质量信号,超过四步后收益递减。
原文摘要 · Abstract (English)
Evaluation of medical artificial intelligence agents remains predominantly answer-centric, assessing only the correctness of final outputs while overlooking the quality of intermediate reasoning. In clinical settings, however, a correct answer reached through fabricated evidence or incoherent logic is as dangerous as an incorrect one. We propose MedTraj, a framework that treats reasoning trajectories as critical objects for construction, evaluation, and optimization. The pipeline generates structured multi-step reasoning chains from medical reasoning sources. Each trajectory is then parsed into clinical observations, evidence, numbered reasoning steps, and a final conclusion, and scored across five quality dimensions: coherence, evidence support, hallucination, completeness, and traceability. Controlled error injection introduces targeted faults into otherwise correct trajectories to establish causal links between specific reasoning failures and measurable quality degradation. Building on this, step-level filtering based on marginal contribution identifies which individual reasoning steps drive or undermine trajectory quality. Finally, quality-weighted context learning feeds trajectory evaluations back into the model at inference time, allowing it to learn from both strong and weak reasoning demonstrations. Experiments across CareQA, PubMedQA, and CECMed demonstrate that trajectory context consistently improves reasoning coherence, with gains of +0.029 to +0.041 over a zero-shot baseline. On CECMed, quality-weighted context nearly doubles the correctness over the zero-shot baseline while cutting the hallucination ratio by 87%. Marginal-contribution analysis further shows that a small minority of reasoning steps carry most of the quality signal, and that extending chains beyond four steps yields diminishing returns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。