用对话式训练提升医学大模型的临床推理能力
Dialogue is Better Than Monologue: Instructing Medical LLMs via Strategical Conversations
- 将静态问答转为模拟真实诊疗的对话形式进行微调
- 在多轮推理中准确率提升9.64%,噪声环境下提升6.18%
- 适合希望提升医疗AI临床实用性的研究者与开发者
当前医疗AI系统在真实临床推理中表现不佳,因其主要基于静态文本和问答任务进行训练与评估,忽视了循证推理与干扰信息处理等关键能力。为此,我们提出一个新基准,模拟真实诊断场景,融入符合USMLE标准的噪声与难度分级。同时,探索基于对话的微调方法,将静态数据转化为对话格式,更贴近迭代式推理过程。实验表明,对话微调模型在多轮推理场景中准确率提升9.64%,在噪声环境中准确率提升6.18%。结果表明,对话微调是构建更符合临床需求且鲁棒的医疗AI系统的有效路径。
原文摘要 · Abstract (English)
Current medical AI systems often fail to replicate real-world clinical reasoning, as they are predominantly trained and evaluated on static text and question-answer tasks. These tuning methods and benchmarks overlook critical aspects like evidence-based reasoning and handling distracting information. To bridge this gap, we introduce a novel benchmark that simulates real-world diagnostic scenarios, integrating noise and difficulty levels aligned with USMLE standards. Moreover, we explore dialogue-based fine-tuning, which transforms static datasets into conversational formats to better capture iterative reasoning processes. Experiments show that dialogue-tuned models outperform traditional methods, with improvements of $9.64\%$ in multi-round reasoning scenarios and $6.18\%$ in accuracy in a noisy environment. Our findings highlight dialogue tuning as a promising approach for advancing clinically aligned and robust medical AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。