把第二视角图像当语言用,提升X光违禁品检测准确率
Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection
- 将双视角图像视为语言模态,融合几何与语义信息进行推理
- 在45,613对双视图数据上,模型任务准确率显著提升
- 适合安全检测、多视角视觉理解等场景的从业者参考
自动X光违禁品检测对安检至关重要,传统方法依赖单视图视觉信息,难以应对复杂威胁。尽管近期研究引入语言引导单视图检测,但实际安检中人类通常使用双视图图像。本文提出首个综合性基准DualXrayBench,包含多视图与多模态数据,支持八项跨视图推理任务。构建了涵盖12类物品的45,613对双视图图像及其对应描述语料。基于此,提出几何-语义交叉模态推理器(GSR),通过结构化思维链<top>、<side>、<conclusion>,联合学习跨视图几何对应与跨模态语义关系,将第二视角图像视为‘语言式’约束。在DualXrayBench上的全面评估表明,GSR在所有任务中均取得显著性能提升,为真实世界X光安检提供新思路。
原文摘要 · Abstract (English)
Automatic X-ray prohibited items detection is vital for security inspection and has been widely studied. Traditional methods rely on visual modality, often struggling with complex threats. While recent studies incorporate language to guide single-view images, human inspectors typically use dual-view images in practice. This raises the question: can the second view provide constraints similar to a language modality? In this work, we introduce DualXrayBench, the first comprehensive benchmark for X-ray inspection that includes multiple views and modalities. It supports eight tasks designed to test cross-view reasoning. In DualXrayBench, we introduce a caption corpus consisting of 45,613 dual-view image pairs across 12 categories with corresponding captions. Building upon these data, we propose the Geometric (cross-view)-Semantic (cross-modality) Reasoner (GSR), a multimodal model that jointly learns correspondences between cross-view geometry and cross-modal semantics, treating the second-view images as a "language-like modality". To enable this, we construct the GSXray dataset, with structured Chain-of-Thought sequences: <top>, <side>, <conclusion>. Comprehensive evaluations on DualXrayBench demonstrate that GSR achieves significant improvements across all X-ray tasks, offering a new perspective for real-world X-ray inspection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。