arXiv:2603.12430cs.CV2026-03被引 2

Surg-R1通过分层推理提升手术决策可解释性,多中心验证表现领先。

Surg-R1: A Hierarchical Reasoning Foundation Model for Scalable and Interpretable Surgical Decision Support with Multi-Center Clinical Validation

  • 构建三层推理架构:感知定位、关系理解、上下文推理
  • 在6个公开与6个多中心数据集上达64.9%胜率,超越主流模型
  • 适合临床医生验证的可解释手术辅助系统,推动精准外科发展

手术场景理解不仅需要准确预测,还需医生可验证的可解释推理。现有视觉语言模型缺乏推理链条,通用推理模型因缺乏领域知识无法胜任复杂手术任务。我们提出Surg-R1,一种通过四阶段训练流程实现分层推理的外科视觉语言模型。核心贡献包括:(1) 三层次推理框架,将手术理解分解为感知定位、关系理解与上下文推理;(2) 构建最大规模的外科思维链数据集,含32万条推理对;(3) 四阶段训练流程,涵盖监督微调、组相对策略优化与迭代自提升。在SurgBench评估中,包含六个公开基准与来自五个机构的六个多中心外部验证数据集,Surg-R1在公开基准上取得64.9%的最高竞技得分,优于Gemini 3.0 Pro(46.1%)与GPT-5.1(37.9%),在器械定位、三元组识别、阶段识别、动作识别及安全视图评估等任务上普遍领先,外部验证中相较最强基线提升15.2个百分点。

原文摘要 · Abstract (English)

Surgical scene understanding demands not only accurate predictions but also interpretable reasoning that surgeons can verify against clinical expertise. However, existing surgical vision-language models generate predictions without reasoning chains, and general-purpose reasoning models fail on compositional surgical tasks without domain-specific knowledge. We present Surg-R1, a surgical Vision-Language Model that addresses this gap through hierarchical reasoning trained via a four-stage pipeline. Our approach introduces three key contributions: (1) a three-level reasoning hierarchy decomposing surgical interpretation into perceptual grounding, relational understanding, and contextual reasoning; (2) the largest surgical chain-of-thought dataset with 320,000 reasoning pairs; and (3) a four-stage training pipeline progressing from supervised fine-tuning to group relative policy optimization and iterative self-improvement. Evaluation on SurgBench, comprising six public benchmarks and six multi-center external validation datasets from five institutions, demonstrates that Surg-R1 achieves the highest Arena Score (64.9%) on public benchmarks versus Gemini 3.0 Pro (46.1%) and GPT-5.1 (37.9%), outperforming both proprietary reasoning models and specialized surgical VLMs on the majority of tasks spanning instrument localization, triplet recognition, phase recognition, action recognition, and critical view of safety assessment, with a 15.2 percentage point improvement over the strongest surgical baseline on external validation.

手术智能可解释性视觉语言模型多中心验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。