arXiv:2605.23068cs.CV2026-05

构建手术视觉问答基准,让AI理解模糊术野中的临床问题。

RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering

论文配图:RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
图 1 · 摘自论文原文
  • 复用公开分割数据集,统一格式生成临床相关问题与答案
  • 覆盖手术流程、解剖结构、图像质量等6类问题,支持精准评估
  • 适合医疗AI研究者,尤其关注手术视觉理解与多模态交互

机器人辅助微创手术中可靠的视觉理解不仅需要精确的掩码,还需应对临床实际中由遮挡、烟雾、出血和反光导致的图像退化。我们提出 extbf{RoboSurg-VQA},一个基于公共手术分割数据集并采用统一标注范式的分割感知视觉问答基准。每帧图像配有一组固定且临床驱动的问题,涵盖手术流程、解剖结构(含区域)、成像模态/视角、手术伪影、图像质量及基本可见性与空间属性,答案为封闭式选项以确保评估一致性。为实现规模化标注,采用约束提示生成候选答案,并通过自动有效性与一致性检查,再经人工审核提升合理性与标签一致性。报告了基准统计数据、基准性能及在复杂手术条件下的常见评估挑战。代码将发布于https://github.com/ziyangwang007/Robosurg-VQA。

原文摘要 · Abstract (English)

Reliable visual understanding in robot-assisted and minimally invasive surgery (RMIS/MIS) demands more than accurate masks: in clinical practice, clinicians pose language-like questions about procedural context, visibility, artefacts, and the presence of anatomical structures and surgical instruments, often under degraded views caused by occlusion, smoke, bleeding, and specular highlights. We present \textbf{RoboSurg-VQA}, a segmentation-aware visual question answering (VQA) benchmark built by repurposing public surgical segmentation datasets under a shared schema. Each frame is paired with a fixed set of clinically motivated questions spanning procedure context, anatomy (including region), imaging modality/view, surgical artefacts, image quality, and basic visibility and spatial attributes, with closed answer sets to enable consistent evaluation. To scale annotation, we generate candidate answers via constrained prompting with automatic validity and consistency checks, followed by human auditing to improve plausibility and label consistency. We report benchmark statistics, sanity baselines, and common evaluation challenges under challenging surgical conditions. The code will be available on https://github.com/ziyangwang007/Robosurg-VQA.

视觉问答手术智能多模态医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。