arXiv:2606.15049cs.CV2026-06

用解剖空间先验提升手术视频中细微结构的检测准确率

Gaussian Spatial Priors for Anatomy-Aware Object Detection in Surgical Videos

  • 在DETR解码器中注入高斯空间先验,利用解剖结构间的位置约束
  • 对小血管等难检结构检测提升33.5%(AP50),显著优于基线模型
  • 适合需要精准定位解剖结构的智能手术辅助系统使用

在腹股沟疝修补术中,检测解剖结构对术中安全至关重要。尽管标准方法可可靠识别柯珀韧带、死亡三角等大结构,但如腹壁上血管等小型结构因视觉模糊和间歇可见而难以检测。我们发现结构间的空间关系具有解剖学约束性,提出高斯空间先验(GSP)模块,将这种约束编码为紧凑参数偏置,注入DAB-DETR解码器的自注意力机制中。该先验通过训练标注离线计算,以一组冻结的高斯参数形式存储,并在每层解码器中基于迭代优化的参考点重新计算。在包含5折交叉验证的腹股沟疝修补术视频数据集上,GSP使依赖类检测的AP50提升33.5%(相比DAB-DETR)和53.9%(相比YOLOv26),同时锚点检测提升6.0%,所有折叠结果均具统计显著性(p=0.012,配对t检验)。

原文摘要 · Abstract (English)

Detecting anatomical structures in surgical video is essential for intraoperative safety frameworks such as the Critical View of Myopectineal Orifice (CVMPO) in inguinal hernia repair. While prominent structures like the Cooper's Ligament and Triangle of Doom are reliably detected by standard methods, smaller structures such as the epigastric vessels remain challenging due to their visual ambiguity and intermittent visibility. We observe that the spatial relationship between structures is anatomically constrained, and propose a Gaussian Spatial Prior (GSP) module that encodes this relationship as a compact, parametric bias injected into the self-attention of a DAB-DETR decoder. The prior is computed offline from training annotations as a small set of frozen Gaussian parameters and recomputed at each decoder layer using the iteratively refined reference points. On a dataset of inguinal hernia repair videos with 5-fold cross-validation, GSP improves dependent class detection by $+33.5\%$ ($\text{AP}_{50}$) over DAB-DETR and $+53.9\%$ over YOLOv26, while also improving anchor detection by $+6.0\%$. These gains are statistically significant across all folds ($p=0.012$, paired $t-$test).

医学图像目标检测解剖先验视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。