让医学大模型像医生一样看图思考,精准定位病灶并推理诊断。
Ophiuchus: Incentivizing Tool-augmented "Think with Images" for Joint Medical Segmentation, Understanding and Reasoning
- 通过三阶段训练让大模型自主决定何时、何地调用视觉工具
- 在多任务医学基准上超越现有主流方法,尤其在分割和推理任务中表现突出
- 适合需要精细图像分析与逻辑推理的医疗AI研究者使用
近期医学多模态大模型在生成逐步文本推理链方面取得显著进展,但仍难以应对需动态、迭代关注细粒度视觉区域的复杂临床任务。为此,我们提出Ophiuchus,一种多功能、工具增强型框架,使医学大模型能够(i)判断何时需要细粒度视觉证据,(ii)确定在医学图像中的探测与定位位置,(iii)将相关子图像内容无缝融入交错的多模态思维链中,实现精准分割与诊断。Ophiuchus不仅调用工具,更将大模型固有的定位与推理能力与外部工具深度融合,提升决策准确性与可信度。核心方法为三阶段训练策略:冷启动监督微调以掌握基础工具选择;自反思微调强化决策修正;代理式工具强化学习激发复杂、专家级诊断行为。大量实验表明,Ophiuchus在多个医学基准(包括VQA、检测与基于推理的分割)上持续优于封闭源与开源的最先进方法。项目代码已开源:https://github.com/SII-zyj/Ophiuchus。
原文摘要 · Abstract (English)
Recent medical MLLMs have made significant progress in generating step-by-step textual reasoning chains. However, they still struggle with complex clinical tasks that necessitate dynamic and iterative focusing on fine-grained visual regions. To close this gap, we introduce Ophiuchus, a versatile, tool-augmented framework that equips an MLLM to (i) decide when fine-grained visual evidence is needed, (ii) determine where to probe and ground within the medical image, and (iii) seamlessly weave the relevant sub-image content back into an interleaved, multimodal chain of thought for precise segmentation and diagnosis. Ophiuchus moves beyond mere tool-calling by tightly fusing the MLLM's inherent grounding and reasoning capabilities with external tools, enabling more accurate and trustworthy decisions. The core of our method is a three-stage training strategy: cold-start SFT for basic tool selection; self-reflection fine-tuning to strengthen decision revision; and agentic tool reinforcement learning to elicit sophisticated, expert-like diagnostic behaviors. Extensive experiments show that Ophiuchus consistently outperforms both closed-source and open-source SOTA methods across diverse medical benchmarks, including VQA, detection, and reasoning-based segmentation. Our project code is available at https://github.com/SII-zyj/Ophiuchus.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。