用导师-实习生协作搜索法,生成高质量医学推理数据
Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative Search
- 导师逐步引导,实习生沿路径推理,动态筛选最优解
- 在多个医疗视觉问答任务上超越现有模型,最高提升12.3%
- 适合医学AI研究者和需要可解释诊断系统的开发者
多模态大语言模型(MLLMs)在通用任务中展现出强大的推理能力,但在医疗领域仍处于起步阶段。构建思维链(CoT)训练数据对增强医疗MLLM的推理能力至关重要。然而,现有方法缺乏全面的框架来搜索和评估关键诊断的有效推理路径。为此,我们提出导师-实习生协作搜索(MICS)机制:先由导师模型逐步初始化推理路径,再由多个实习生模型沿该路径继续思考,最后根据多实习生模型的综合推理表现选择最优路径。推理质量由MICS-Score评估。最终,我们构建了多任务医疗推理数据集MMRP(按难度排序),并基于课程学习策略训练出新模型Chiron-o1,具备强大的视觉问答与泛化推理能力。大量实验表明,使用MICS构建的CoT数据训练的Chiron-o1,在多项医疗视觉问答与推理基准测试中达到当前最优性能。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have begun to demonstrate robust reasoning capabilities on general tasks, yet their application in the medical domain remains in its early stages. Constructing chain-of-thought (CoT) training data is essential for bolstering the reasoning abilities of medical MLLMs. However, existing approaches exhibit a deficiency in offering a comprehensive framework for searching and evaluating effective reasoning paths towards critical diagnosis. To address this challenge, we propose Mentor-Intern Collaborative Search (MICS), a novel reasoning-path searching scheme to generate rigorous and effective medical CoT data. MICS first leverages mentor models to initialize the reasoning, one step at a time, then prompts each intern model to continue the thinking along those initiated paths, and finally selects the optimal reasoning path according to the overall reasoning performance of multiple intern models. The reasoning performance is determined by an MICS-Score, which assesses the quality of generated reasoning paths. Eventually, we construct MMRP, a multi-task medical reasoning dataset with ranked difficulty, and Chiron-o1, a new medical MLLM devised via a curriculum learning strategy, with robust visual question-answering and generalizable reasoning capabilities. Extensive experiments demonstrate that Chiron-o1, trained on our CoT dataset constructed using MICS, achieves state-of-the-art performance across a list of medical visual question answering and reasoning benchmarks. Codes are available at https://github.com/manglu097/Chiron-o1
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。