用双阶段模型实时定位内镜手术中出血点,提升止血效率。
BleedOrigin: Dynamic Bleeding Source Localization in Endoscopic Submucosal Dissection via Dual-Stage Detection and Tracking
- 分两阶段检测与追踪,从出血发生到持续定位全流程处理
- 在106,222帧上实现96.85%的出血起始检测准确率
- 首个专用数据集,覆盖8个解剖部位和6类复杂场景
内镜黏膜下剥离术(ESD)中术中出血带来重大风险,需精准、实时定位出血源并持续监测以实现有效止血。由于内镜医生需反复冲洗清除血液,仅能抓住毫秒级时机识别出血点,效率低且延长手术时间,增加患者风险。现有AI方法多聚焦于出血区域分割,忽视了在频繁视觉遮挡与动态变化的复杂环境中对出血源的精准检测与时间追踪。这一问题因缺乏专用数据集而加剧。为此,我们构建了首个全面的ESD出血源数据集BleedOrigin-Bench,包含44例手术中106,222帧上的1,771个专家标注出血源,另补充39,755帧伪标签。该数据集覆盖8个解剖部位及6类临床挑战场景。同时提出BleedOrigin-Net,一种新型双阶段检测-追踪框架,实现从出血起始检测到空间连续追踪的完整流程。对比YOLOv11/v12、多模态大语言模型及点追踪方法,实验表明其达到96.85%帧级准确率(±≤8帧),70.24%像素级初始定位准确率(≤100像素),96.11%像素级追踪准确率(≤100像素),性能领先。
原文摘要 · Abstract (English)
Intraoperative bleeding during Endoscopic Submucosal Dissection (ESD) poses significant risks, demanding precise, real-time localization and continuous monitoring of the bleeding source for effective hemostatic intervention. In particular, endoscopists have to repeatedly flush to clear blood, allowing only milliseconds to identify bleeding sources, an inefficient process that prolongs operations and elevates patient risks. However, current Artificial Intelligence (AI) methods primarily focus on bleeding region segmentation, overlooking the critical need for accurate bleeding source detection and temporal tracking in the challenging ESD environment, which is marked by frequent visual obstructions and dynamic scene changes. This gap is widened by the lack of specialized datasets, hindering the development of robust AI-assisted guidance systems. To address these challenges, we introduce BleedOrigin-Bench, the first comprehensive ESD bleeding source dataset, featuring 1,771 expert-annotated bleeding sources across 106,222 frames from 44 procedures, supplemented with 39,755 pseudo-labeled frames. This benchmark covers 8 anatomical sites and 6 challenging clinical scenarios. We also present BleedOrigin-Net, a novel dual-stage detection-tracking framework for the bleeding source localization in ESD procedures, addressing the complete workflow from bleeding onset detection to continuous spatial tracking. We compare with widely-used object detection models (YOLOv11/v12), multimodal large language models, and point tracking methods. Extensive evaluation demonstrates state-of-the-art performance, achieving 96.85% frame-level accuracy ($\pm\leq8$ frames) for bleeding onset detection, 70.24% pixel-level accuracy ($\leq100$ px) for initial source detection, and 96.11% pixel-level accuracy ($\leq100$ px) for point tracking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。