arXiv:2507.04270cs.CVcs.AI2025-07被引 1

ZERO模型通过多模态提示实现工业场景零样本部署,无需重训练即可通用。

ZERO: Industry-ready Vision Foundation Model with Multi-modal Prompts

  • 用文本和图像双模态提示增强泛化能力,避免重新训练。
  • 在37个工业数据集上表现优于现有模型,且在两项国际竞赛中分别获第2、第4名。
  • 专为工业领域设计,适合数据少、需快速落地的现实应用。

基础模型虽已革新AI,但在真实工业场景中因缺乏高质量领域特定数据,难以实现零样本部署。为此,Superb AI推出了面向工业场景的视觉基础模型ZERO,通过文本与视觉多模态提示实现无需重训练的泛化能力。该模型在自有千亿级工业数据集中精选0.9百万标注样本进行训练,在LVIS-Val等学术基准上表现优异,并在37个多样化的工业数据集上显著超越现有模型。此外,ZERO在CVPR 2025目标实例检测挑战赛中获第2名,在基础少样本目标检测挑战赛中获第4名,充分证明其在少量适配和有限数据下具备出色的实用性与泛化能力。据我们所知,ZERO是首个专为领域特定零样本工业应用构建的视觉基础模型。

原文摘要 · Abstract (English)

Foundation models have revolutionized AI, yet they struggle with zero-shot deployment in real-world industrial settings due to a lack of high-quality, domain-specific datasets. To bridge this gap, Superb AI introduces ZERO, an industry-ready vision foundation model that leverages multi-modal prompting (textual and visual) for generalization without retraining. Trained on a compact yet representative 0.9 million annotated samples from a proprietary billion-scale industrial dataset, ZERO demonstrates competitive performance on academic benchmarks like LVIS-Val and significantly outperforms existing models across 37 diverse industrial datasets. Furthermore, ZERO achieved 2nd place in the CVPR 2025 Object Instance Detection Challenge and 4th place in the Foundational Few-shot Object Detection Challenge, highlighting its practical deployability and generalizability with minimal adaptation and limited data. To the best of our knowledge, ZERO is the first vision foundation model explicitly built for domain-specific, zero-shot industrial applications.

视觉基础模型工业应用多模态提示零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。