无需训练,用逻辑与局部搜索结合实现可追溯的工业缺陷检测。
Global Logic and Local Search: Dual-Stream Multimodal In-Context Learning for Verifiable Industrial Anomaly Detection

- 构建视觉-逻辑地图,整合正常样本与规范信息作为参考。
- 全局流提取可验证视觉事实,局部流在预算内搜索关键区域证据。
- 适合缺乏缺陷样本的早期部署场景,决策全程可追溯。
大型多模态模型虽具备强少样本泛化能力,但工业缺陷检测仍面临挑战:缺陷尺寸小、输入分辨率有限,且文本标准常与视觉证据脱节。现有基于优化的方法虽能提升对齐效果,但通常需大量缺陷样本,难以在早期部署阶段应用。本文提出无需训练的GLLS框架,实现参考引导的多模态上下文验证。GLLS利用部件感知的视觉-逻辑图谱,在推理上下文中组织正常参考与结构化规范。其融合全局逻辑流(由SAM 3提取部分可验证视觉事实)与细粒度动作流(在固定预算内通过MCTS选择局部证据区域)。在MMAD-QA及其他异常检测数据集上的实验表明,该方法持续优于匹配基线与通用基线,同时保证诊断决策始终可追溯至明确的视觉证据。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) show strong few-shot generalization, but industrial anomaly detection remains difficult because defects are small, input resolution is limited, and textual standards are not always grounded in visual evidence. Recent optimization-based methods improve alignment through fine-tuning, but they often require many defective samples, which are unavailable in early deployment. We present Global Logic and Local Search (GLLS), a training-free framework for reference-guided multimodal in-context verification. GLLS uses a Part-Aware Visual-Logical Atlas to organize normal references and structured specifications in the inference context. It combines a Global & Logic Stream, where SAM 3 extracts partially checkable visual facts, with a Fine-Grained & Actions Stream, where MCTS selects local evidence crops under a fixed budget. Experiments on MMAD-QA and additional anomaly detection datasets show consistent gains over matched and general-purpose baselines, while keeping the final diagnostic decision traceable to explicit visual evidence throughout the inspection trace.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。