评测手术视觉语言模型在器械与操作识别上的表现,揭示其依赖弱上下文而非临床视觉证据。
SurgXBench: Explainable Vision-Language Model Benchmark for Surgery
- 基于双公开数据集测试多个先进视觉语言模型的零样本性能
- 发现模型多依赖弱上下文线索,而非真实手术视觉特征
- 引入可解释AI分析注意力机制,提供预测可靠性评估新视角
数字智能创新正推动机器人手术向更智能决策发展,实时感知手术器械存在与操作(如切割组织)至关重要。然而,尽管研究数十年,多数机器学习模型仍基于小规模数据集,泛化能力不足。近期视觉语言模型(VLMs)在跨模态推理上取得突破性进展,其卓越泛化能力预示着在智能机器人手术中的巨大潜力。但手术领域对VLMs的研究仍不充分,现有模型表现有限,亟需基准测试以评估其能力与局限,指导未来研发。为此,我们评估了多个先进VLMs在两个公开的机器人辅助腹腔镜数据集上进行器械与动作分类的零样本性能。除常规评估外,还融合可解释AI技术可视化模型注意力,并挖掘预测背后的因果解释,为该领域提供前所未有的可靠性评估视角。我们提出若干基于可解释性分析的度量指标,以补充传统评估方式。分析表明,尽管经过领域特定训练,手术VLMs常依赖弱上下文线索而非临床相关视觉证据,凸显在手术应用中加强视觉与推理监督的必要性。
原文摘要 · Abstract (English)
Innovations in digital intelligence are transforming robotic surgery with more informed decision-making. Real-time awareness of surgical instrument presence and actions (e.g., cutting tissue) is essential for such systems. Yet, despite decades of research, most machine learning models for this task are trained on small datasets and still struggle to generalize. Recently, vision-Language Models (VLMs) have brought transformative advances in reasoning across visual and textual modalities. Their unprecedented generalization capabilities suggest great potential for advancing intelligent robotic surgery. However, surgical VLMs remain under-explored, and existing models show limited performance, highlighting the need for benchmark studies to assess their capabilities and limitations and to inform future development. To this end, we benchmark the zero-shot performance of several advanced VLMs on two public robotic-assisted laparoscopic datasets for instrument and action classification. Beyond standard evaluation, we integrate explainable AI to visualize VLM attention and uncover causal explanations behind their predictions. This provides a previously underexplored perspective in this field for evaluating the reliability of model predictions. We also propose several explainability analysis-based metrics to complement standard evaluations. Our analysis reveals that surgical VLMs, despite domain-specific training, often rely on weak contextual cues rather than clinically relevant visual evidence, highlighting the need for stronger visual and reasoning supervision in surgical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。