arXiv:2411.17217cs.CV2024-11AAAI被引 12

让SAM更懂工业异常,通过自我感知调优提升分割精度

Promptable Anomaly Segmentation with SAM Through Self-Perception Tuning

  • 用自草稿调优策略生成初始异常掩码并迭代优化
  • 在多个基准数据集上显著超越基线方法,提升明显
  • 适合需要高精度异常检测的工业视觉场景

Segment Anything Model (SAM) 凭借出色的泛化能力在异常分割任务中取得显著进展。然而,现有通过提示直接应用 SAM 的方法常忽略领域偏移问题——SAM 在自然图像上表现良好,但在工业场景中性能下降。参数高效微调(PEFT)虽具潜力,却因未能充分应对异常图像中的感知挑战而表现欠佳。本文提出一种新型自我感知调优(SPT)方法,旨在增强 SAM 对异常分割的感知能力。SPT 采用自草稿调优策略,先生成初始粗略异常掩码,再进行细化;同时引入视觉关系感知适配器,提升对判别性关系信息的感知以优化掩码生成。大量实验证明,SPT 在多个基准数据集上显著优于基线方法,验证了其有效性。

原文摘要 · Abstract (English)

Segment Anything Model (SAM) has made great progress in anomaly segmentation tasks due to its impressive generalization ability. However, existing methods that directly apply SAM through prompting often overlook the domain shift issue, where SAM performs well on natural images but struggles in industrial scenarios. Parameter-Efficient Fine-Tuning (PEFT) offers a promising solution, but it may yield suboptimal performance by not adequately addressing the perception challenges during adaptation to anomaly images. In this paper, we propose a novel \textbf{S}elf-\textbf{P}erceptinon \textbf{T}uning (\textbf{SPT}) method, aiming to enhance SAM's perception capability for anomaly segmentation. The SPT method incorporates a self-drafting tuning strategy, which generates an initial coarse draft of the anomaly mask, followed by a refinement process. Additionally, a visual-relation-aware adapter is introduced to improve the perception of discriminative relational information for mask generation. Extensive experimental results on several benchmark datasets demonstrate that our SPT method can significantly outperform baseline methods, validating its effectiveness.

异常分割SAM视觉感知工业检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。