用统一视觉语言模型提升X光安检识别能力,支持多任务且泛化性强。
OneFocus: Enabling Real-World X-ray Security Screening with a Unified Vision-Language Model

- 构建统一视觉语言模型OneFocus,支持问答、定位、分类与理解四类任务。
- 在52,124张图像-文本对上训练,实现跨域泛化,性能达当前最优。
- 适合安全筛查、智能物流等需快速识别违禁品的现实场景使用。
X光违禁品检测对大规模物流与交通安保至关重要,但传统检测器难以适应新型违禁品且缺乏基础视觉理解能力。视觉语言模型(VLMs)虽具强泛化性,却受限于高质量X光图像-描述数据稀缺。为此,我们提出MMXray,一个精心构建的基准数据集,包含52,124对图像-文本,覆盖28个细粒度违禁品类别。为引入真实遮挡模式,我们进一步构建CleanDET合成数据集,包含28类清洁前景违禁品图像及多种密度背景图像,并设计AnyContraSyn可控合成方法。同时开发OnePipe可扩展的数据整理管道。基于MMXray,我们提出OneFocus——一个统一的VLM,支持视觉问答、违禁品定位、分类与图像理解四项核心任务。OneFocus在X光违禁品理解上达到当前最佳性能,并展现出强大跨域泛化能力,为安全筛查建立坚实视觉语言基线。
原文摘要 · Abstract (English)
X-ray contraband detection is critical for security in large-scale logistics and transportation, yet conventional detectors struggle to adapt to emerging contraband types and lack fundamental visual understanding. Vision-language models (VLMs) offer strong generalization but are hindered by the scarcity of high-quality X-ray image-caption data. To bridge this critical gap, we present MMXray, a meticulously curated benchmark of 52,124 image-caption pairs spanning 28 fine-grained classes of X-ray contraband. To enrich MMXray with realistic occlusion patterns, we further introduce CleanDET, a dedicated synthesis dataset containing clean foreground contraband images from 28 categories and background images with diverse density levels, together with AnyContraSyn, a controllable synthesis method designed to operate on CleanDET. We also develop OnePipe, an extensible pipeline for systematic data curation. Built on MMXray, we propose OneFocus, a unified VLM that supports four core tasks: visual question answering, contraband localization, classification, and image understanding. OneFocus achieves state-of-the-art performance in X-ray contraband understanding and demonstrates robust cross-domain generalization, establishing a strong vision-language baseline for security screening.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。