arXiv:2604.20319cs.CV2026-04中稿 · CVPR

构建手术视频时空推理新基准,评估大模型理解手术细节能力

SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark

论文配图:SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark
图 1 · 摘自论文原文
  • 设计包含5类推理维度的链式思考框架,用结构化标注提升评估精度
  • 测试10个主流多模态大模型,发现商业模型表现最优但整体仍存明显差距
  • 适用于医疗AI研究者,助力提升手术视频分析的逻辑推理能力

对手术视频进行细粒度时空推理至关重要,但多模态大语言模型(MLLMs)在此领域的表现尚未充分探索。为填补这一空白,我们提出SurgCoT,一个统一的基准,用于评估7个外科专业、35种不同手术中MLLMs的链式思考(CoT)推理能力。SurgCoT通过结构化的CoT框架(问题-选项-知识-线索-答案),评估五项核心推理维度:因果动作排序、线索-动作对齐、可操作性映射、微过渡定位和异常起始追踪。其中,'知识'字段提供必要背景信息,'线索'字段给出明确的时空证据。对10个领先MLLMs的评估显示:1)商业模型优于开源及医学专用版本;2)手术场景下链式思考推理仍存在显著差距;3)SurgCoT能有效评估并促进渐进式时空推理能力提升。该基准为缩小大模型能力与临床推理需求之间的鸿沟提供了可复现的测试环境。代码已开源:https://github.com/CVI-SZU/SurgCoT。

原文摘要 · Abstract (English)

Fine-grained spatiotemporal reasoning on surgical videos is critical, yet the capabilities of Multi-modal Large Language Models (MLLMs) in this domain remain largely unexplored. To bridge this gap, we introduce SurgCoT, a unified benchmark for evaluating chain-of-thought (CoT) reasoning in MLLMs across 7 surgical specialties and 35 diverse procedures. SurgCoT assesses five core reasoning dimensions: Causal Action Ordering, Cue-Action Alignment, Affordance Mapping, Micro-Transition Localization, and Anomaly Onset Tracking, through a structured CoT framework with an intensive annotation protocol (Question-Option-Knowledge-Clue-Answer), where the Knowledge field provides essential background context and Clue provides definitive spatiotemporal evidence. Evaluation of 10 leading MLLMs shows: 1) commercial models outperform open-source and medical-specialized variants; 2) significant gaps exist in surgical CoT reasoning; 3) SurgCoT enables effective evaluation and enhances progressive spatiotemporal reasoning. SurgCoT provides a reproducible testbed to narrow the gap between MLLM capabilities and clinical reasoning demands. Code: https://github.com/CVI-SZU/SurgCoT.

手术视频链式思考多模态推理评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。