AI可像人一样用双视角X光图识别违禁品,性能提升超24%
Dual-view X-ray Detection: Can AI Detect Prohibited Items from Dual-view X-ray Images like Humans?
- 用主视角与辅视角协同检测,模仿人类双视判断逻辑
- 在伞类等难题上检测率提升24.7%,整体性能显著增强
- 适配多种检测模型,适用于安检AI系统研发
为检测高难度类别中的违禁物品,人类检查员通常依赖垂直与侧向两个视角的图像。人工智能能否以类似方式利用双视角X光图像进行检测?现有X光数据集常受限于单视角成像或样本多样性不足。为此,我们提出了大规模双视角X光数据集LDXray,涵盖12个类别共353,646个实例,为模型训练与评估提供丰富资源。为模拟人类双视角智能,我们提出辅助视角增强网络(AENet),该框架同时利用同一物体的主视角与辅助视角。主视角分支专注于常见类别检测,而辅助视角分支通过主视角学习的“专家模型”处理更难类别。在LDXray数据集上的大量实验表明,双视角机制显著提升检测性能,例如伞类挑战性类别最高提升24.7%。此外,AENet在七种不同检测模型中均表现出强泛化能力。
原文摘要 · Abstract (English)
To detect prohibited items in challenging categories, human inspectors typically rely on images from two distinct views (vertical and side). Can AI detect prohibited items from dual-view X-ray images in the same way humans do? Existing X-ray datasets often suffer from limitations, such as single-view imaging or insufficient sample diversity. To address these gaps, we introduce the Large-scale Dual-view X-ray (LDXray), which consists of 353,646 instances across 12 categories, providing a diverse and comprehensive resource for training and evaluating models. To emulate human intelligence in dual-view detection, we propose the Auxiliary-view Enhanced Network (AENet), a novel detection framework that leverages both the main and auxiliary views of the same object. The main-view pipeline focuses on detecting common categories, while the auxiliary-view pipeline handles more challenging categories using ``expert models" learned from the main view. Extensive experiments on the LDXray dataset demonstrate that the dual-view mechanism significantly enhances detection performance, e.g., achieving improvements of up to 24.7% for the challenging category of umbrellas. Furthermore, our results show that AENet exhibits strong generalization across seven different detection models for X-ray Inspection
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。