用手术讲座语音构建大规模推理数据集,提升AI对术中决策的理解能力。
SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning
- 从手术教学视频中自动提取语音讲解,构建12类问题的结构化问答数据
- 生成206.8万组问答对,在354个专家验证样本上达84%准确率
- 提出视觉语言模型与强化学习推理模型,可推断操作意图与风险
外科医生不仅观察,更会解读:专家在手术场景中不仅能识别器械,还能理解其选择原因、潜在风险及后续步骤。现有外科AI无法回答此类问题,主要因缺乏标注明确手术推理的数据。然而,手术教学视频已包含这些解释——专家为教学目的所作的意图、理由与预测叙述。尽管原始内容嘈杂无序,却蕴含当前外科AI所缺失的推理信号。我们提出SUREON,一个大规模视频问答数据集,系统性地从外科学术视频中挖掘训练信号。SUREON定义了涵盖安全评估、决策理由与预测的12类问题,采用多智能体流水线实现规模化提取与结构化标注。覆盖134.7万个片段与170种术式,生成206.8万个问答对,并建立由专家验证的354例基准测试集。为评估该监督信号对推理能力的转化效果,我们引入两个模型:SureonVLM(经监督微调的视觉语言模型)与SureonVLM-R1(基于组相对策略优化训练的推理模型)。两者均能回答复杂手术问题,显著优于更大规模的通用领域模型,在SUREON基准上准确率超84%,且在标准外科感知任务中表现更优。对SureonVLM-R1的定性分析显示其具备显式的推理行为,如从视觉上下文推断手术意图。
原文摘要 · Abstract (English)
Surgeons don't just see -- they interpret. When an expert observes a surgical scene, they understand not only what instrument is being used, but why it was chosen, what risk it poses, and what comes next. Current surgical AI cannot answer such questions, largely because training data that explicitly encodes surgical reasoning is immensely difficult to annotate at scale. Yet surgical video lectures already contain exactly this -- explanations of intent, rationale, and anticipation, narrated by experts for the purpose of teaching. Though inherently noisy and unstructured, these narrations encode the reasoning that surgical AI currently lacks. We introduce SUREON, a large-scale video QA dataset that systematically harvests this training signal from surgical academic videos. SUREON defines 12 question categories covering safety assessment, decision rationale, and forecasting, and uses a multi-agent pipeline to extract and structure supervision at scale. Across 134.7K clips and 170 procedure types, SUREON yields 206.8k QA pairs and an expert-validated benchmark of 354 examples. To evaluate the extent to which this supervision translates to surgical reasoning ability, we introduce two models: SureonVLM, a vision-language model adapted through supervised fine-tuning, and SureonVLM-R1, a reasoning model trained with Group Relative Policy Optimization. Both models can answer complex questions about surgery and substantially outperform larger general-domain models, exceeding 84% accuracy on the SUREON benchmark while outperforming general-domain models on standard surgical perception tasks. Qualitative analysis of SureonVLM-R1 reveals explicit reasoning behavior, such as inferring operative intent from visual context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。