arXiv:2608.14262cs.CV2026-08中稿 · MICCAI 2026

测试外科内镜视频中视觉语言模型在常见图像退化下的鲁棒性,提出新评测基准并改进模型表现。

On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos

论文配图:On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos
图 1 · 摘自论文原文
  • 构建针对内镜视频的六类真实退化扰动数据集Endo-C6,固定高严重度评估
  • 发现现有模型在极端退化下性能显著下降,最差情况准确率暴跌至20%以下
  • 通过轻量微调提升模型鲁棒性,无需改动提示接口,适合临床部署

时序视觉语言模型(TVLMs)为外科视频理解提供了可复用的提示式接口,但其在内镜实际采集中的退化干扰下鲁棒性尚未充分评估。实践中,模糊、雾霾、运动模糊、噪声、电凝烟雾和数据包丢失等退化引入结构化分布偏移,可能破坏视频与文本对齐。本文研究了由片段帧退化引发的时序VLM鲁棒性,提出端到端-6(Endo-C6),一个包含六种内镜真实扰动的紧凑基准,以固定高严重度评估公共胃肠道(GI)内镜和腹腔镜胆囊切除术视频。在标准化提示协议下,对3个近期外科TVLM基线进行294次数据集级评估,分析其在均值与最差情况下的表现。最终提出RobustEndoCLIP,通过VeRA实现少量样本参数高效微调,优于现有TVLM基线。结果表明,现成的TVLM在特定内镜退化下可能出现严重最差情况崩溃,而轻量少样本适应可显著提升退化场景下的性能与鲁棒性,且不改变提示接口。我们期望Endo-C6推动标准化鲁棒性报告,促进更可靠的临床视觉语言系统发展。

原文摘要 · Abstract (English)

Temporal vision-language models (TVLMs) offer a reusable, prompt-based interface for surgical video understanding, yet, their robustness under clinically realistic acquisition artifacts in endoscopy remains insufficiently characterized. In practice, degradations such as defocus, haze, motion blur, noise, cautery smoke, and packet loss introduce structured distribution shifts which may compromise video-text alignment. We study the robustness of temporal VLMs under such shifts caused by corruptions in clip frames. We introduce Endo-C6, a compact corruption benchmark of six endoscopy-realistic perturbations evaluated at a fixed high severity, and apply it to public Gastrointestinal (GI) endoscopy and laparoscopic cholecystectomy videos. Under a standardized prompt protocol, we benchmark 3 recent surgical TVLM baselines and analyze robustness in both mean and worst-case settings, spanning 294 dataset-level evaluations. Finally, we present RobustEndoCLIP, obtained by few-shot parameter-efficient tuning with VeRA, outperforming existing TVLM baselines. Our findings show that off-the-shelf TVLMs can exhibit severe worst-case collapse under endoscopy-specific corruptions, whereas lightweight few-shot adaptation can substantially improve corrupted performance and robustness without changing the prompt-based interface. We expect Endo-C6 to support standardized robustness reporting and promote more reliable clinical vision-language systems.

视觉语言模型手术视频鲁棒性评测内镜分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。