arXiv:2501.17860cs.CLcs.AI2025-01被引 2

用对话式训练提升医学大模型的临床推理能力

Dialogue is Better Than Monologue: Instructing Medical LLMs via Strategical Conversations

  • 将静态问答转为模拟真实诊疗的对话形式进行微调
  • 在多轮推理中准确率提升9.64%,噪声环境下提升6.18%
  • 适合希望提升医疗AI临床实用性的研究者与开发者

当前医疗AI系统在真实临床推理中表现不佳,因其主要基于静态文本和问答任务进行训练与评估,忽视了循证推理与干扰信息处理等关键能力。为此,我们提出一个新基准,模拟真实诊断场景,融入符合USMLE标准的噪声与难度分级。同时,探索基于对话的微调方法,将静态数据转化为对话格式,更贴近迭代式推理过程。实验表明,对话微调模型在多轮推理场景中准确率提升9.64%,在噪声环境中准确率提升6.18%。结果表明,对话微调是构建更符合临床需求且鲁棒的医疗AI系统的有效路径。

原文摘要 · Abstract (English)

Current medical AI systems often fail to replicate real-world clinical reasoning, as they are predominantly trained and evaluated on static text and question-answer tasks. These tuning methods and benchmarks overlook critical aspects like evidence-based reasoning and handling distracting information. To bridge this gap, we introduce a novel benchmark that simulates real-world diagnostic scenarios, integrating noise and difficulty levels aligned with USMLE standards. Moreover, we explore dialogue-based fine-tuning, which transforms static datasets into conversational formats to better capture iterative reasoning processes. Experiments show that dialogue-tuned models outperform traditional methods, with improvements of $9.64\%$ in multi-round reasoning scenarios and $6.18\%$ in accuracy in a noisy environment. Our findings highlight dialogue tuning as a promising approach for advancing clinically aligned and robust medical AI systems.

医学大模型对话微调临床推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。