arXiv:2506.19469cs.CVcs.AI2025-06被引 6

用强化学习训练的多模态大模型,让手术视觉问答更懂推理。

Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning

  • 构建手术场景的思维链数据集,分两阶段微调模型提升推理能力。
  • 在手术视觉问答任务上超越现有最先进模型,准确率显著提升。
  • 适合医疗AI研究者和手术机器人开发者参考,推动临床应用落地。

近年来,手术场景理解取得显著进展,尤其在机器人手术中的视觉问题定位回答(Surgical-VQLA)任务上。然而,现有模型缺乏深层推理能力和可解释性,限制了其在临床中的可靠性与发展潜力。为此,受推理型多模态大模型启发,我们构建了首个包含视觉问答、定位问答及思维链(CoT)配对数据的Surgery-R1-54k数据集。提出首个专用于Surgical-VQLA的推理型多模态大模型(Surgery-R1),采用监督微调(SFT)与强化微调(RFT)两阶段机制,赋予基础模型复杂推理能力。为实现高效高质量的奖励机制,设计了多模态一致性奖励,缓解手术场景中可能存在的位置错觉问题。实验表明,Surgery-R1在Surgical-VQLA任务上优于其他先进模型,并验证了其推理能力与方法有效性。代码与数据集将公开于https://github.com/FiFi-HAO467/Surgery-R1。

原文摘要 · Abstract (English)

In recent years, significant progress has been made in the field of surgical scene understanding, particularly in the task of Visual Question Localized-Answering in robotic surgery (Surgical-VQLA). However, existing Surgical-VQLA models lack deep reasoning capabilities and interpretability in surgical scenes, which limits their reliability and potential for development in clinical applications. To address this issue, inspired by the development of Reasoning Multimodal Large Language Models (MLLMs), we first build the Surgery-R1-54k dataset, including paired data for Visual-QA, Grounding-QA, and Chain-of-Thought (CoT). Then, we propose the first Reasoning MLLM for Surgical-VQLA (Surgery-R1). In our Surgery-R1, we design a two-stage fine-tuning mechanism to enable the basic MLLM with complex reasoning abilities by utilizing supervised fine-tuning (SFT) and reinforcement fine-tuning (RFT). Furthermore, for an efficient and high-quality rule-based reward system in our RFT, we design a Multimodal Coherence reward mechanism to mitigate positional illusions that may arise in surgical scenarios. Experiment results demonstrate that Surgery-R1 outperforms other existing state-of-the-art (SOTA) models in the Surgical-VQLA task and widely-used MLLMs, while also validating its reasoning capabilities and the effectiveness of our approach. The code and dataset will be organized in https://github.com/FiFi-HAO467/Surgery-R1.

手术AI多模态推理模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。