arXiv:2606.27084cs.CVeess.IV2026-06

用伪文本提示实现腹部CT器官精确定位,轻量高效且开源可用。

Pseudo-Text-Conditioned 3D Grounding DINO for Organ Localization in Abdominal CT

论文配图:Pseudo-Text-Conditioned 3D Grounding DINO for Organ Localization in Abdominal CT
图 1 · 摘自论文原文
  • 用冻结的伪文本令牌替代真实文本编码器,实现3D器官定位。
  • 在193个数据上达到0.5830的mAP(IoU 0.1~0.7),粗略定位准确率高达0.9649。
  • 适合医学影像分析、创伤评估与多模态预训练研究者参考。

腹部CT中可靠的器官定位可为后续创伤分析提供空间先验。我们提出CT-3GDINO,一种轻量级3D检测器,采用类似Grounding-DINO的基于查询的架构,通过冻结的伪文本类别标记实现固定器官定位,无需真实文本编码器。模型结合Swin3D视觉主干、双向特征增强、伪文本引导的查询选择和跨模态解码器,预测肝脏、脾脏、左肾、右肾和肠管的归一化3D边界框。在193个匹配的RSNA/RATIC CT体积上训练与评估,使用分割生成的标注框。最佳多尺度模型从头训练,于3D IoU阈值0.1至0.7范围内取得0.5830的整体顶1类别mAP,优于固定主干(0.5570)与可训练主干(0.4657)变体。粗略定位表现优异(IoU=0.1时AP=0.9649),但严格对齐能力有限(IoU=0.7时AP=0.1552)。结果确立了CT-3GDINO作为伪文本条件3D器官定位的开源基线,并推动未来在定位感知预训练、更丰富多模态条件及损伤导向检测方面的研究。

原文摘要 · Abstract (English)

Reliable organ localization in abdominal CT can provide spatial priors for downstream trauma analysis. We propose CT-3GDINO, a lightweight 3D detector that adapts a Grounding-DINO-style query-based architecture to fixed organ localization using frozen pseudo-text class tokens instead of a real text encoder. The model combines a Swin3D visual backbone, bidirectional feature enhancement, pseudo-text-guided query selection, and a cross-modality decoder to predict normalized 3D boxes for liver, spleen, left kidney, right kidney, and bowel. We train and evaluate on 193 matched RSNA/RATIC CT volumes with segmentation-derived boxes. The best multi-scale model, trained from scratch, achieves 0.5830 overall top-1 class-wise mAP over 3D IoU thresholds from 0.1 to 0.7, outperforming fixed- and trainable-backbone classification-pretrained variants with 0.5570 and 0.4657 mAP. Performance is strong for coarse localization, with 0.9649 AP at IoU 0.1, but remains limited for strict box alignment, with 0.1552 AP at IoU 0.7. These results establish CT-3GDINO as an open-source baseline for pseudo-text-conditioned 3D organ localization and motivate future work on localization-aware pretraining, richer multimodal conditioning, and injury-focused detection.

3D定位医学影像伪文本CT分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。