通过合成数据与测试时训练,提升大模型医学推理能力。
FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training
- 用高质量医学合成数据进行监督微调和偏好优化
- 在医疗基准上性能比之前模型平均提升23%
- 首次引入测试时训练,适合医疗推理场景的模型优化
近年来,大语言模型在疾病诊断和治疗方案制定等医疗应用中展现出潜力。然而,大多数现有医疗大模型在应对复杂医学问题(如鉴别诊断和用药建议)时仍缺乏深度推理能力。本文提出FineMedLM-o1,利用高质量医学合成数据和长文本推理数据,通过监督微调(SFT)与直接偏好优化(DPO)实现高级对话与深层推理能力。此外,首次在医疗领域引入测试时训练(TTT),实现领域自适应并保障推理可靠性。实验表明,FineMedLM-o1在关键医疗基准上相比先前模型平均性能提升23%。进一步引入TTT带来额外14%的性能增益,验证其在增强医学推理方面的有效性。为支持该过程,我们还提出一种新颖的医学对话生成方法。相较其他开源数据集,本研究构建的数据集在质量和复杂度上均更优。项目代码与数据将公开于GitHub。
原文摘要 · Abstract (English)
Recent advancements in large language models (LLMs) have shown promise in medical applications such as disease diagnosis and treatment planning. However, most existing medical LLMs struggle with the deep reasoning required for complex medical problems, such as differential diagnosis and medication recommendations. We propose FineMedLM-o1, which leverages high-quality medical synthetic data and long-form reasoning data for Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO), enabling advanced dialogue and deep reasoning capabilities. Additionally, we introduce Test-Time Training (TTT) in the medical domain for the first time, facilitating domain adaptation and ensuring reliable, accurate reasoning. Experimental results demonstrate that FineMedLM-o1 achieves a 23% average performance improvement over prior models on key medical benchmarks. Furthermore, the introduction of TTT provides an additional 14% performance boost, highlighting its effectiveness in enhancing medical reasoning capabilities. To support this process, we also propose a novel method for synthesizing medical dialogue. Compared to other open-source datasets, our dataset stands out as superior in both quality and complexity. The project and data will be released on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。