arXiv:2503.17071cs.CV2025-03ICCV被引 6

无需训练,用现成模型实现精准的X光物品检测。

Superpowering Open-Vocabulary Object Detectors for X-ray Vision

论文配图:Superpowering Open-Vocabulary Object Detectors for X-ray Vision
图 1 · 摘自论文原文
  • 用网页图像和材料迁移生成高质量X光类特征
  • 相比基础模型平均提升17.0点mAP
  • 适合安全筛查与开放词汇检测研究者

开放词汇目标检测(OvOD)有望革新安检流程,使系统能识别X光扫描中的任意物品。然而,由于数据稀缺和模态差异,直接应用基于RGB的解决方案面临挑战。为此,我们提出RAXO——一种无需训练的框架,可复用现成的RGB OvOD检测器实现鲁棒的X光检测。RAXO采用双源检索策略构建高质量的X光类别描述符:从网络获取相关RGB图像,并通过创新的X光材料迁移机制进行增强,避免依赖标注数据库。这些视觉描述符替代了传统文本分类,利用模态内特征距离实现检测。大量实验表明,RAXO在多个基准上持续提升性能,相较基线模型平均提升17.0点mAP。为进一步推动该领域发展,我们还发布了DET-COMPASS,一个包含超过300个物体类别的边界框标注新基准,支持大规模评估X光场景下的OvOD。代码与数据集见:https://github.com/PAGF188/RAXO。

原文摘要 · Abstract (English)

Open-vocabulary object detection (OvOD) is set to revolutionize security screening by enabling systems to recognize any item in X-ray scans. However, developing effective OvOD models for X-ray imaging presents unique challenges due to data scarcity and the modality gap that prevents direct adoption of RGB-based solutions. To overcome these limitations, we propose RAXO, a training-free framework that repurposes off-the-shelf RGB OvOD detectors for robust X-ray detection. RAXO builds high-quality X-ray class descriptors using a dual-source retrieval strategy. It gathers relevant RGB images from the web and enriches them via a novel X-ray material transfer mechanism, eliminating the need for labeled databases. These visual descriptors replace text-based classification in OvOD, leveraging intra-modal feature distances for robust detection. Extensive experiments demonstrate that RAXO consistently improves OvOD performance, providing an average mAP increase of up to 17.0 points over base detectors. To further support research in this emerging field, we also introduce DET-COMPASS, a new benchmark featuring bounding box annotations for over 300 object categories, enabling large-scale evaluation of OvOD in X-ray. Code and dataset available at: https://github.com/PAGF188/RAXO.

开放词汇检测X光分析零样本检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。