构建首个真实场景X光行李安检多模态数据集,训练出能理解指令的视觉AI助手。
STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspection
- 基于真实扫描构建图文配对数据集,支持多模态指令理解
- 在21类威胁上实现跨域泛化,性能优于现有方法
- 适合安全检测、AI辅助审检等实际应用场景
计算机辅助筛查(CAS)系统的发展对提升X光行李扫描中的安全威胁检测至关重要。然而,当前数据集难以反映真实世界中复杂威胁与隐藏手段,且现有方法受限于预定义标签的闭集范式。为此,我们提出STCray,首个多模态X光行李安检数据集,包含46,642张图像-文本配对扫描,覆盖21类威胁类别,使用机场安检用X射线扫描仪生成。该数据集采用专门设计的协议,确保领域感知、连贯的文本描述,形成可用于多模态指令遵循的高质量数据。基于此,我们训练了名为STING-BEE的领域感知视觉语言模型,支持场景理解、指代定位、视觉定位和视觉问答等任务,为X光行李安检中的多模态学习建立了新基准。此外,STING-BEE在跨域设置下表现出最先进的泛化能力。代码、数据与模型已公开于https://divs1159.github.io/STING-BEE/。
原文摘要 · Abstract (English)
Advancements in Computer-Aided Screening (CAS) systems are essential for improving the detection of security threats in X-ray baggage scans. However, current datasets are limited in representing real-world, sophisticated threats and concealment tactics, and existing approaches are constrained by a closed-set paradigm with predefined labels. To address these challenges, we introduce STCray, the first multimodal X-ray baggage security dataset, comprising 46,642 image-caption paired scans across 21 threat categories, generated using an X-ray scanner for airport security. STCray is meticulously developed with our specialized protocol that ensures domain-aware, coherent captions, that lead to the multi-modal instruction following data in X-ray baggage security. This allows us to train a domain-aware visual AI assistant named STING-BEE that supports a range of vision-language tasks, including scene comprehension, referring threat localization, visual grounding, and visual question answering (VQA), establishing novel baselines for multi-modal learning in X-ray baggage security. Further, STING-BEE shows state-of-the-art generalization in cross-domain settings. Code, data, and models are available at https://divs1159.github.io/STING-BEE/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。