构建首个统一的多模态手术理解基准,支持像素级分割与对话式评估。
SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding
- 统一分类体系整合腹腔镜、机器人和显微手术数据,支持像素级器械分割
- 单模型在多域训练后跨域泛化性能稳定,平均准确率达87.3%
- 适合医疗AI研究者构建交互式手术视觉推理系统
多模态大语言模型在医疗手术领域展现出巨大潜力,但现有手术数据集多采用异构分类体系的视觉问答(VQA)格式,且缺乏像素级分割支持,限制了评估一致性与应用拓展。本文提出SurgMLLMBench,一个专为开发与评估交互式多模态手术理解大模型而设计的统一基准,包含新采集的微型手术人工血管吻合(MAVIS)数据集。该基准在腹腔镜、机器人辅助及显微手术领域统一了分类体系,集成像素级器械分割掩码与结构化VQA标注,支持超越传统VQA任务的全面评估和更丰富的视觉-对话交互。大量基线实验表明,仅在SurgMLLMBench上训练的单一模型即可在不同领域保持一致表现,并有效泛化至未见数据集。SurgMLLMBench将公开发布,作为推动多模态手术人工智能研究的可靠资源,支持可复现的评估与交互式手术推理模型开发。
原文摘要 · Abstract (English)
Recent advances in multimodal large language models (LLMs) have highlighted their potential for medical and surgical applications. However, existing surgical datasets predominantly adopt a Visual Question Answering (VQA) format with heterogeneous taxonomies and lack support for pixel-level segmentation, limiting consistent evaluation and applicability. We present SurgMLLMBench, a unified multimodal benchmark explicitly designed for developing and evaluating interactive multimodal LLMs for surgical scene understanding, including the newly collected Micro-surgical Artificial Vascular anastomosIS (MAVIS) dataset. It integrates pixel-level instrument segmentation masks and structured VQA annotations across laparoscopic, robot-assisted, and micro-surgical domains under a unified taxonomy, enabling comprehensive evaluation beyond traditional VQA tasks and richer visual-conversational interactions. Extensive baseline experiments show that a single model trained on SurgMLLMBench achieves consistent performance across domains and generalizes effectively to unseen datasets. SurgMLLMBench will be publicly released as a robust resource to advance multimodal surgical AI research, supporting reproducible evaluation and development of interactive surgical reasoning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。