构建首个面向微小物体的多视角图文异常检测数据集
MANTA: A Large-Scale Multi-View and Visual-Text Anomaly Detection Dataset for Tiny Objects
- 采集38类物体137.3万张多视角图像,含8.6万张像素级标注异常图
- 包含875个描述异常的术语及2000道多选题,覆盖异常成因与视觉特征
- 适用于跨模态异常检测研究,尤其适合微小目标场景
我们提出MANTA,一个面向微小物体的视觉-文本异常检测数据集。视觉部分包含跨越五个典型领域的38个物体类别,共超过137.3万张图像,其中8.6万张被标注为异常,并提供像素级注释。每张图像从五个不同视角拍摄,确保物体全覆盖。文本部分包含两个子集:声明性知识(Declarative Knowledge)涵盖875个描述各类域中常见异常的词汇,包含<什么、为何、如何>的详细解释,包括原因与视觉特征;建构性学习(Constructivist Learning)提供2000道难度各异的多选题,每题配图并附答案解析。我们还提出了视觉-文本任务的基线方法,并进行了广泛的基准测试,评估多种先进方法在不同设置下的表现,凸显了该数据集的挑战性与有效性。
原文摘要 · Abstract (English)
We present MANTA, a visual-text anomaly detection dataset for tiny objects. The visual component comprises over 137.3K images across 38 object categories spanning five typical domains, of which 8.6K images are labeled as anomalous with pixel-level annotations. Each image is captured from five distinct viewpoints to ensure comprehensive object coverage. The text component consists of two subsets: Declarative Knowledge, including 875 words that describe common anomalies across various domains and specific categories, with detailed explanations for < what, why, how>, including causes and visual characteristics; and Constructivist Learning, providing 2K multiple-choice questions with varying levels of difficulty, each paired with images and corresponded answer explanations. We also propose a baseline for visual-text tasks and conduct extensive benchmarking experiments to evaluate advanced methods across different settings, highlighting the challenges and efficacy of our dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。