用弱监督方法实现工业农业中目标与异常的精准分割
PASTA: Vision Transformer Patch Aggregation for Weakly Supervised Target and Anomaly Segmentation
- 通过对比场景与标准参考,利用ViT特征空间分布分析定位目标与异常
- 训练时间减少75.8%,工业和农业场景下目标与异常分割准确率分别达88.3%和63.5%IoU
- 无需精细标注,适合缺乏标注数据的工业质检与农作除草场景
在材料回收和除草等非结构化环境中检测未见异常是一项关键挑战。现有感知系统因依赖大量人工标注数据,难以满足实时处理、像素级分割精度和鲁棒性要求。为此,我们提出一种弱监督目标与异常分割流水线PASTA,仅需图像级弱标签。PASTA通过对比观测场景与标准参考,在自监督视觉变压器(ViT)特征空间中进行分布分析,识别目标与异常。该方法结合段落任何模型3(SAM3)的语义文本提示,实现零样本对象分割。在自建的钢铁废料回收数据集和植物数据集上的评估显示,相比领域特定基线,本方法训练时间减少75.8%。尽管具备领域无关性,仍可在工业与农业场景中达到最高88.3% IoU的目标分割性能和最高63.5% IoU的异常分割性能。
原文摘要 · Abstract (English)
Detecting unseen anomalies in unstructured environments presents a critical challenge for industrial and agricultural applications such as material recycling and weeding. Existing perception systems frequently fail to satisfy the strict operational requirements of these domains, specifically real-time processing, pixel-level segmentation precision, and robust accuracy, due to their reliance on exhaustively annotated datasets. To address these limitations, we propose a weakly supervised pipeline for object segmentation and classification using weak image-level supervision called 'Patch Aggregation for Segmentation of Targets and Anomalies' (PASTA). By comparing an observed scene with a nominal reference, PASTA identifies Target and Anomaly objects through distribution analysis in self-supervised Vision Transformer (ViT) feature spaces. Our pipeline utilizes semantic text-prompts via the Segment Anything Model 3 to guide zero-shot object segmentation. Evaluations on a custom steel scrap recycling dataset and a plant dataset demonstrate a 75.8% training time reduction of our approach to domain-specific baselines. While being domain-agnostic, our method achieves superior Target (up to 88.3% IoU) and Anomaly (up to 63.5% IoU) segmentation performance in the industrial and agricultural domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。