用800万条带结构化推理轨迹的数据,训练出更强的医疗多模态模型。
OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
- 构建含结构化推理轨迹的数据集,指导模型逐步推理解题。
- 在68亿个响应标记上训练,多个医学基准测试达到开源模型顶尖水平。
- 模型能自主调节推理长度,适应不同临床任务,无需额外标注。
高质量且精心策划的数据是训练医疗大语言模型的基础,直接影响模型的泛化能力和对未见临床任务的鲁棒性。本文研究了训练策略与数据筛选方法,以构建医学领域内具备强健多模态推理能力的模型。重点聚焦于监督微调(SFT),探索利用结构化推理轨迹的数据配方。通过所提出的配方,我们扩展实验至超过800万条样本、68亿个响应标记的数据集,使模型在多种分布外的医学基准任务中表现优于现有开源模型。结果进一步表明,构建高质量、多样化的训练数据集,并包含不同长度的结构化推理轨迹,可使微调后的模型在无显式监督下,根据下游任务自动校准其推理路径长度。本文还分享关键洞见,描述数据构建策略,并提出未来构建鲁棒医疗视觉-语言推理系统的关键方向。
原文摘要 · Abstract (English)
High-quality and carefully curated data is a cornerstone of training medical large language models, as it directly impacts both generalization and robustness to unseen clinical tasks. We investigate strategies for training and data curation to develop a robust multimodal reasoning model in the medical domain. Our work focuses on supervised fine-tuning (SFT) and explores data recipes that leverage structured reasoning traces. Using our proposed data recipe, we scale experiments to a dataset of over 8 million examples and 6.8 billion response tokens, achieving state-of-the-art performance among open-source models across diverse out-of-distribution medical benchmark tasks. Our results further indicate that curating a high-quality, diverse training dataset with varying structured reasoning trace lengths enables the fine-tuned model to self-calibrate its reasoning trajectory lengths based on the downstream task, without explicit supervision. We present key insights, describe the data curation strategy, and outline next steps toward developing robust medical vision-language reasoning system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。