arXiv:2409.12741cs.CLcs.AI2024-09被引 8

医学大模型微调中,直接偏好优化提升复杂任务表现。

Fine Tuning Large Language Models for Medicine: The Role and Importance of Direct Preference Optimization

  • 对比SFT与DPO在医学文本任务中的表现
  • DPO显著提升临床推理等复杂任务效果
  • 揭示当前工具链对DPO支持不足

大型语言模型(LLM)在医学领域的微调应用仍不充分。常见的微调方法包括监督微调(SFT)和直接偏好优化(DPO),但缺乏使用指导。本研究比较了SFT与DPO在五类常见医学自然语言任务中的表现:基于文本的分类、基于数值的分类、临床推理、摘要生成和临床分诊。结果表明,仅用SFT即可满足基于文本的分类任务,而临床推理、摘要生成和临床分诊等更复杂任务中,DPO能显著提升性能。研究确立了DPO在医学领域的作用与重要性,并指出当前软件工具链存在缺陷,阻碍该技术的广泛应用。

原文摘要 · Abstract (English)

Large Language Model (LLM) fine tuning is underutilized in the field of medicine. Two of the most common methods of fine tuning are Supervised Fine Tuning (SFT) and Direct Preference Optimization (DPO), but there is little guidance informing users when to use either technique. In this investigation, we compare the performance of SFT and DPO for five common natural language tasks in medicine: Classification with text data, Classification with numeric data, Clinical Reasoning, Summarization, and Clinical Triage. We find that SFT alone is sufficient for Classification with text data, whereas DPO improves performance for the more complex tasks of Clinical Reasoning, Summarization and Clinical Triage. Our results establish the role and importance of DPO fine tuning within medicine, and consequently call attention to current software gaps that prevent widespread deployment of this technique.

大模型微调医学AIDPO临床推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。