现有视觉语言模型对缺陷检测的文本控制能力被高估,新基准揭示其实际效果有限。
A Structured Benchmark for Text-Guided Anomaly Detection: When Language Stops Conditioning the Decision

- 设计三阶段文本引导异常检测基准,逐步提升语言作用
- 多模型测试显示文本仅浅层影响决策,去除关键词后性能骤降
- 适合关注工业检测可解释性与语言可控性的研究者
工业异常检测传统上是单模态任务。近年来,多模态视觉-语言模型虽支持文本输入,声称实现零样本和少样本检测,但评估仍沿用单模态基准,固定文本条件,无法衡量语言是否真正影响决策;所谓性能提升可能源于强预训练视觉特征而非文本引导。本文提出文本引导异常检测(TGAD)结构化基准,包含三个渐进场景:在MVTec AD上的受控提示敏感性测试;扩展后的组件标注版MVTec AD,要求模型仅关注指定部件;以及新的组装面板数据集(APD),模拟真实工业场景,需同时识别缺陷类型和部件位置。评估三类代表性模型:生成式大视觉-语言、免训练判别式、嵌入自适应判别式。结果表明,文本接口仅表面影响判断:移除物体名词后,生成模型的I-AUROC从97.4降至82.6;当缺陷超出指令区域仍被视作正常时,性能由90.3跌至66.3;在APD上两者结合时,图像级判别性能低于MVTec水平,部分低于随机水平(71.2, 50.5, 31.5)。结果表明,现有基准夸大了当前多模态异常检测系统的文本引导能力,此类评估协议是工业部署中可靠语言控制的前提。
原文摘要 · Abstract (English)
Industrial anomaly detection has historically been a unimodal task. Recent multimodal vision-language models have produced systems that admit textual input alongside the image and are presented as enabling text-guided zero- and few-shot inspection. Yet these methods are evaluated with protocols inherited from unimodal benchmarks that hold the textual condition constant and therefore cannot measure whether language conditions the decision; whether reported gains reflect text guidance or strong pretrained visual features remains open. We introduce Text-Guided Anomaly Detection (TGAD), a structured benchmark that progressively increases the functional role of language across three scenarios: a controlled prompt-sensitivity setting on MVTec AD; a component-tagged extension of MVTec AD that requires the model to restrict its assessment to an instructed part; and the new Assembled Panel Dataset (APD), a realistic industrial setting that requires both defect-type and component-location knowledge. We evaluate one representative model per paradigm: generative large vision-language, training-free discriminative, and embedding-adaptive discriminative. In all three, the textual interface conditions the decision only superficially: prompt content is absorbed unless the object noun is removed (the generative model's I-AUROC drops from 97.4 to 82.6); component-level instructions do not constrain the decision once defects outside the instructed part are admitted as normal (from 90.3 to 66.3); and when both combine on APD, image-level discrimination collapses below the MVTec level, in one case below chance (71.2, 50.5, 31.5). These results suggest that standard benchmarks overstate the text-guided capabilities of current multimodal anomaly detection systems, and that a protocol of this kind is a prerequisite for models that can be reliably controlled through language for industrial deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。