arXiv:2506.07196cs.CVcs.CL2025-06被引 1

构建手术动作规划评测基准,评估大模型在真实手术场景中的推理能力。

SAP-Bench: Benchmarking Multimodal Large Language Models in Surgical Action Planning

  • 基于胆囊切除术构建带时间标注的动作数据集,含1226个临床验证片段
  • 7大顶尖多模态模型在该任务上表现不足,平均准确率未达理想水平
  • 适合医疗AI、手术机器人研发者,推动可解释性决策系统发展

有效评估对推动多模态大模型(MLLM)研究至关重要。手术动作规划(SAP)任务旨在从视觉输入生成未来动作序列,需具备精确且复杂的分析能力。与数学推理不同,手术决策涉及生命安全,要求严谨可验证的流程以保障可靠性与患者安全。当前基准难以评估模型区分原子动作与协调长时程复杂操作的能力。为此,我们提出SAP-Bench,一个大规模高质量数据集,用于支持多模态大模型进行可解释的手术动作规划。该基准基于胆囊切除术,平均持续时长1137.5秒,包含1,226个临床验证的动作片段(平均时长68.7秒),覆盖五类基础手术动作,来自74个手术过程。数据集提供1,152个策略采样的当前帧,每帧对应下一个动作作为多模态分析锚点。我们提出MLLM-SAP框架,利用多模态大模型从当前手术场景和自然语言指令中生成下一步动作建议,并注入手术领域知识。为评估数据集有效性及模型整体能力,我们评测了七种主流多模态大模型(如OpenAI-o1、GPT-4o、QwenVL2.5-72B、Claude-3.5-Sonnet、GeminiPro2.5、Step-1o、GLM-4v),揭示了当前模型在下一步动作预测上的显著性能差距。

原文摘要 · Abstract (English)

Effective evaluation is critical for driving advancements in MLLM research. The surgical action planning (SAP) task, which aims to generate future action sequences from visual inputs, demands precise and sophisticated analytical capabilities. Unlike mathematical reasoning, surgical decision-making operates in life-critical domains and requires meticulous, verifiable processes to ensure reliability and patient safety. This task demands the ability to distinguish between atomic visual actions and coordinate complex, long-horizon procedures, capabilities that are inadequately evaluated by current benchmarks. To address this gap, we introduce SAP-Bench, a large-scale, high-quality dataset designed to enable multimodal large language models (MLLMs) to perform interpretable surgical action planning. Our SAP-Bench benchmark, derived from the cholecystectomy procedures context with the mean duration of 1137.5s, and introduces temporally-grounded surgical action annotations, comprising the 1,226 clinically validated action clips (mean duration: 68.7s) capturing five fundamental surgical actions across 74 procedures. The dataset provides 1,152 strategically sampled current frames, each paired with the corresponding next action as multimodal analysis anchors. We propose the MLLM-SAP framework that leverages MLLMs to generate next action recommendations from the current surgical scene and natural language instructions, enhanced with injected surgical domain knowledge. To assess our dataset's effectiveness and the broader capabilities of current models, we evaluate seven state-of-the-art MLLMs (e.g., OpenAI-o1, GPT-4o, QwenVL2.5-72B, Claude-3.5-Sonnet, GeminiPro2.5, Step-1o, and GLM-4v) and reveal critical gaps in next action prediction performance.

手术规划多模态模型医疗AI评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。