首个评估神经外科解剖识别能力的多模态基准,揭示大模型仍远低于人类水平。
NeuroABench: A Multimodal Evaluation Benchmark for Neurosurgical Anatomy Identification
- 构建覆盖89个术式的9小时标注视频数据集,评估68个解剖结构识别能力。
- 顶尖大模型仅达40.87%准确率,低于人类平均46.5%和最低人类28%。
- 专为神经外科解剖理解设计,适合医疗AI、医学教育与手术辅助研究者。
多模态大语言模型(MLLMs)在手术视频理解方面展现出巨大潜力,具备更强的零样本能力与人机交互效率,为推进手术教育与辅助提供基础。然而,现有研究与数据集主要聚焦于手术流程与操作理解,对关键的解剖认知关注不足。临床中,外科医生高度依赖精准解剖知识来解读、复盘与学习手术视频。为填补这一空白,我们提出神经外科解剖基准(NeuroABench),首个专为评估神经外科领域解剖理解而设计的多模态基准。NeuroABench包含9小时标注的神经外科视频,涵盖89个不同手术,采用创新的多轮多模态标注流程构建。该基准评估68个临床解剖结构的识别能力,提供严谨标准化的评估框架。对超过10个前沿MLLM的实验显示显著局限性,最佳模型仅达40.87%准确率。为进一步验证,我们抽取子集并邀请4名神经外科住院医师测试,结果显示最佳学生达56%准确率,最低28%,平均46.5%。尽管最优模型表现接近最低人类水平,但仍显著落后于群体平均水平,凸显了当前模型在解剖理解上的进步与与人类水平之间的巨大差距。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have shown significant potential in surgical video understanding. With improved zero-shot performance and more effective human-machine interaction, they provide a strong foundation for advancing surgical education and assistance. However, existing research and datasets primarily focus on understanding surgical procedures and workflows, while paying limited attention to the critical role of anatomical comprehension. In clinical practice, surgeons rely heavily on precise anatomical understanding to interpret, review, and learn from surgical videos. To fill this gap, we introduce the Neurosurgical Anatomy Benchmark (NeuroABench), the first multimodal benchmark explicitly created to evaluate anatomical understanding in the neurosurgical domain. NeuroABench consists of 9 hours of annotated neurosurgical videos covering 89 distinct procedures and is developed using a novel multimodal annotation pipeline with multiple review cycles. The benchmark evaluates the identification of 68 clinical anatomical structures, providing a rigorous and standardized framework for assessing model performance. Experiments on over 10 state-of-the-art MLLMs reveal significant limitations, with the best-performing model achieving only 40.87% accuracy in anatomical identification tasks. To further evaluate the benchmark, we extract a subset of the dataset and conduct an informative test with four neurosurgical trainees. The results show that the best-performing student achieves 56% accuracy, with the lowest scores of 28% and an average score of 46.5%. While the best MLLM performs comparably to the lowest-scoring student, it still lags significantly behind the group's average performance. This comparison underscores both the progress of MLLMs in anatomical understanding and the substantial gap that remains in achieving human-level performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。