arXiv:2507.07939cs.CL2025-07中稿 · ACMMM2025被引 10

SAGE通过增强事实与熵感知对齐,提升工业异常检测的可解释性与泛化能力。

SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware Alignment

  • 用自引导事实增强融合领域知识,改进视觉推理
  • 在零样本和单样本设置下超越现有方法,准确率显著提升
  • 适合需要可解释异常分析的工业场景应用

尽管视觉语言模型(VLMs)在通用多模态任务中表现优异,但在工业异常检测与推理中常面临可解释性差、难以泛化到未见类别等问题。这源于异常检测本身具有强领域特性,限制了现有VLM在需精确、结构化与上下文感知分析的工业场景中的应用。为此,我们提出SAGE,一种基于VLM的框架,通过自引导事实增强(SFE)与熵感知直接偏好优化(E-DPO)提升异常推理能力。SFE通过事实抽取与融合将领域知识融入视觉推理,E-DPO则利用熵感知优化对齐模型输出与专家偏好。此外,我们构建了AD-PL数据集,包含28,415个问答样本及专家打分响应,专用于工业异常推理。为评估模型逻辑一致性,我们设计了多尺度逻辑评估(MLE)框架。SAGE在零样本与单样本设置下于工业异常检测数据集上表现卓越。代码、模型与数据集已开源。

原文摘要 · Abstract (English)

While Vision-Language Models (VLMs) have shown promising progress in general multimodal tasks, they often struggle in industrial anomaly detection and reasoning, particularly in delivering interpretable explanations and generalizing to unseen categories. This limitation stems from the inherently domain-specific nature of anomaly detection, which hinders the applicability of existing VLMs in industrial scenarios that require precise, structured, and context-aware analysis. To address these challenges, we propose SAGE, a VLM-based framework that enhances anomaly reasoning through Self-Guided Fact Enhancement (SFE) and Entropy-aware Direct Preference Optimization (E-DPO). SFE integrates domain-specific knowledge into visual reasoning via fact extraction and fusion, while E-DPO aligns model outputs with expert preferences using entropy-aware optimization. Additionally, we introduce AD-PL, a preference-optimized dataset tailored for industrial anomaly reasoning, consisting of 28,415 question-answering instances with expert-ranked responses. To evaluate anomaly reasoning models, we develop Multiscale Logical Evaluation (MLE), a quantitative framework analyzing model logic and consistency. SAGE demonstrates superior performance on industrial anomaly datasets under zero-shot and one-shot settings. The code, model and dataset are available at https://github.com/amoreZgx1n/SAGE.

异常检测视觉语言模型可解释性工业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。