arXiv:2602.06619cs.CV2026-02

用因果启发的视觉语言模型,让手术视频识别跨域更准。

CauCLIP: Bridging the Sim-to-Real Gap in Surgical Video Understanding via Causality-Inspired Vision-Language Modeling

  • 通过频率增强和因果抑制损失,学习跨域不变特征
  • 在SurgVisDom上准确率显著超越现有方法
  • 适合需要跨真实与仿真数据部署的医疗AI研究者

手术阶段识别是智能手术室上下文感知决策支持的关键,但受限于标注临床视频少以及仿真与真实手术数据间的大领域差距。为此,我们提出CauCLIP,一种基于因果启发的视觉语言框架,利用CLIP学习无需目标域数据的域不变表示。该方法结合频域增强策略扰动域特定属性,同时保留语义结构,并引入因果抑制损失以消除非因果偏差、强化因果手术特征。这些组件整合进统一训练框架,使模型聚焦于手术流程中的稳定因果因素。在SurgVisDom硬适应基准上的实验表明,本方法显著优于所有对比方法,验证了因果引导的视觉语言模型在域泛化手术视频理解中的有效性。

原文摘要 · Abstract (English)

Surgical phase recognition is a critical component for context-aware decision support in intelligent operating rooms, yet training robust models is hindered by limited annotated clinical videos and large domain gaps between synthetic and real surgical data. To address this, we propose CauCLIP, a causality-inspired vision-language framework that leverages CLIP to learn domain-invariant representations for surgical phase recognition without access to target domain data. Our approach integrates a frequency-based augmentation strategy to perturb domain-specific attributes while preserving semantic structures, and a causal suppression loss that mitigates non-causal biases and reinforces causal surgical features. These components are combined in a unified training framework that enables the model to focus on stable causal factors underlying surgical workflows. Experiments on the SurgVisDom hard adaptation benchmark demonstrate that our method substantially outperforms all competing approaches, highlighting the effectiveness of causality-guided vision-language models for domain-generalizable surgical video understanding.

手术视频因果建模跨域泛化视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。