用视觉语言模型自动识别白内障手术中三大并发症,提升安全预警能力。
CataractCompDetect: Intraoperative Complication Detection in Cataract Surgery
- 融合手术阶段感知与视觉语言推理,精准定位并发症
- 在首个标注数据集上平均F1达70.63%,关键事件检测准确率超60%
- 适合手术训练系统与智能手术助手研发者参考
白内障手术是全球最常见的外科手术之一,但术中虹膜脱出、后囊破裂(PCR)和玻璃体丢失等并发症仍是导致不良结果的主要原因。自动化检测这些事件可实现早期预警并提供客观培训反馈。本文提出CataractCompDetect框架,结合阶段感知定位、基于SAM 2的跟踪、专用于并发症的风险评分以及视觉-语言推理进行最终分类。为验证该框架,我们构建了首个标注术中并发症的白内障手术视频数据集CataComp,包含53例手术,其中23例存在临床并发症。在CataComp上,CataractCompDetect平均F1得分为70.63%,各并发症表现分别为:虹膜脱出81.8%、后囊破裂60.87%、玻璃体丢失69.23%。结果表明,结合结构化手术先验与视觉-语言推理对识别罕见但高影响的术中事件具有显著价值。论文所用数据集与代码将在录用后公开。
原文摘要 · Abstract (English)
Cataract surgery is one of the most commonly performed surgeries worldwide, yet intraoperative complications such as iris prolapse, posterior capsule rupture (PCR), and vitreous loss remain major causes of adverse outcomes. Automated detection of such events could enable early warning systems and objective training feedback. In this work, we propose CataractCompDetect, a complication detection framework that combines phase-aware localization, SAM 2-based tracking, complication-specific risk scoring, and vision-language reasoning for final classification. To validate CataractCompDetect, we curate CataComp, the first cataract surgery video dataset annotated for intraoperative complications, comprising 53 surgeries, including 23 with clinical complications. On CataComp, CataractCompDetect achieves an average F1 score of 70.63%, with per-complication performance of 81.8% (Iris Prolapse), 60.87% (PCR), and 69.23% (Vitreous Loss). These results highlight the value of combining structured surgical priors with vision-language reasoning for recognizing rare but high-impact intraoperative events. Our dataset and code will be publicly released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。