arXiv:2604.00784cs.CV2026-04被引 1

构建手术视频细粒度时空理解数据集,提升视觉语言模型手术场景理解能力

An Approach to Enriching Surgical Video Datasets for Fine-Grained Spatial-Temporal Understanding of Vision-Language Models

论文配图:An Approach to Enriching Surgical Video Datasets for Fine-Grained Spatial-Temporal Understanding of Vision-Language Models
图 1 · 摘自论文原文
  • 提出SurgSTU-Pipeline,通过时空连续性过滤生成高质量标注数据
  • 创建含150万问答对的SurgSTU数据集,覆盖6711段手术视频片段
  • 验证了微调后模型在手术时空理解任务上表现最优,适合医疗AI研究者

手术视频理解是推进计算机辅助手术的关键前提。尽管视觉语言模型(VLMs)已应用于手术领域,现有手术视觉语言数据集仍难以捕捉和评估复杂的交错式时空动态。由于人工标注成本高或大语言模型生成易出错,构建大规模且准确反映手术视频细粒度时空关系的数据集极具挑战。为此,我们提出SurgSTU-Pipeline,一种基于时间与空间连续性过滤的确定性生成流程,可可靠构建用于细粒度时空多模态理解的手术数据集。将该流程应用于公开手术数据集,我们构建了SurgSTU数据集,包含6711个视频片段,密集扩展为15万条细粒度时空问答样本。全面评估表明,尽管当前主流通用型VLMs在零样本设置下表现不佳,但其时空能力可通过上下文学习提升。在SurgSTU训练数据上微调后的VLM在所有时空任务中达到最佳性能,验证了该数据集对提升手术视频中VLM时空理解的有效性。

原文摘要 · Abstract (English)

Surgical video understanding is a crucial prerequisite for advancing Computer-Assisted Surgery. While vision-language models (VLMs) have recently been applied to the surgical domain, existing surgical vision-language datasets lack in capturing and evaluating complex, interleaved spatial-temporal dynamics. Creating large scale datasets that accurately represent fine-grained spatial-temporal relationships in surgical videos is challenging due to costly manual annotations or error-prone generation using large language models. To address this gap, we introduce the SurgSTU-Pipeline, a deterministic generation pipeline featuring temporal and spatial continuity filtering to reliably create surgical datasets for fine-grained spatial-temporal multimodal understanding. Applying this pipeline to publicly available surgical datasets, we create the SurgSTU dataset, comprising 6711 video clips densely extended with 150k fine-grained spatial-temporal question-answer samples. Our comprehensive evaluation shows that while state-of-the-art generalist VLMs struggle in zero-shot settings, their spatial-temporal capabilities can be improved through in-context learning. A fine-tuned VLM on the SurgSTU training dataset achieves highest performance among all spatial-temporal tasks, validating the dataset's efficacy to improve spatial-temporal understanding of VLMs in surgical videos. The project is available here: https://lennart-maack.github.io/SurgSTU-project

手术视频视觉语言模型时空理解数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。