测试大模型在医疗任务中推理模式的效果,发现提升有限。
Benchmarking the Thinking Mode of Multimodal Large Language Models in Clinical Tasks
- 对比两种大模型在推理与非推理模式下的表现
- 多数医疗任务性能提升微弱,复杂问题仍不理想
- 适合关注医疗AI可靠性与改进方向的研究者
多模态大语言模型(MLLMs)近年发展出可显式控制内部思考过程的“推理模式”(thinking mode),区别于传统的“非推理模式”。随着这类“双状态”模型快速应用,本文系统评估了其在医疗任务中的表现。针对两个主流模型Seed1.5-VL和Gemini-2.5-Flash,我们在VQA-RAD和ROCOv2数据集上测试了四类视觉医学任务。结果表明,激活推理模式后,多数任务性能提升有限;在开放问答和医学图像理解等复杂任务上,模型表现仍不理想,凸显出对领域特定医疗数据及更先进知识融合方法的需求。
原文摘要 · Abstract (English)
A recent advancement in Multimodal Large Language Models (MLLMs) research is the emergence of "reasoning MLLMs" that offer explicit control over their internal thinking processes (normally referred as the "thinking mode") alongside the standard "non-thinking mode". This capability allows these models to engage in a step-by-step process of internal deliberation before generating a final response. With the rapid transition to and adoption of these "dual-state" MLLMs, this work rigorously evaluated how the enhanced reasoning processes of these MLLMs impact model performance and reliability in clinical tasks. This paper evaluates the active "thinking mode" capabilities of two leading MLLMs, Seed1.5-VL and Gemini-2.5-Flash, for medical applications. We assessed their performance on four visual medical tasks using VQA-RAD and ROCOv2 datasets. Our findings reveal that the improvement from activating the thinking mode remains marginal compared to the standard non-thinking mode for the majority of the tasks. Their performance on complex medical tasks such as open-ended VQA and medical image interpretation remains suboptimal, highlighting the need for domain-specific medical data and more advanced methods for medical knowledge integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。