用视觉语言模型+上下文学习,零样本识别太赫兹图像中的隐性威胁
Smart Eyes for Silent Threats: VLMs and In-Context Learning for THz Imaging
- 不微调模型,通过提示工程让大模型理解太赫兹图像
- 零样本和单样本下分类准确率显著优于传统方法
- 适合数据少、标注难的科研场景,如安检与材料检测
太赫兹(THz)成像可用于安全筛查和材料分类等非侵入式分析,但因标注有限、分辨率低和视觉模糊,图像分类仍具挑战。本文提出在无微调情况下,利用视觉语言模型(VLMs)结合上下文学习(ICL),实现灵活且可解释的分类。通过模态对齐的提示框架,我们将两个开源大模型适配至太赫兹领域,并在零样本和单样本设置下进行评估。结果表明,该方法在低数据条件下显著提升分类性能与可解释性。这是首次将ICL增强的VLM应用于太赫兹成像,为资源受限的科学领域提供新方向。代码已公开于GitHub。
原文摘要 · Abstract (English)
Terahertz (THz) imaging enables non-invasive analysis for applications such as security screening and material classification, but effective image classification remains challenging due to limited annotations, low resolution, and visual ambiguity. We introduce In-Context Learning (ICL) with Vision-Language Models (VLMs) as a flexible, interpretable alternative that requires no fine-tuning. Using a modality-aligned prompting framework, we adapt two open-weight VLMs to the THz domain and evaluate them under zero-shot and one-shot settings. Our results show that ICL improves classification and interpretability in low-data regimes. This is the first application of ICL-enhanced VLMs to THz imaging, offering a promising direction for resource-constrained scientific domains. Code: \href{https://github.com/Nicolas-Poggi/Project_THz_Classification/tree/main}{GitHub repository}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。