arXiv:2602.19823cs.CV2026-02

无需训练,用超像素融合实现工业场景的开放词汇3D感知

Open-vocabulary 3D scene perception in industrial environments

  • 用预计算的超像素合并生成语义掩码,不依赖预训练实例分割模型
  • 在工业场景中对常见物体的分割准确率显著优于现有方法
  • 适合需要快速部署、无需标注数据的工业视觉应用

生产、物流或制造环境中的自主视觉应用需具备超越固定类别集合的感知能力。当前基于2D视觉-语言基础模型(VLFM)的开放词汇方法通常依赖非工业数据集(如家庭场景)预训练的类无关分割模型,但这些模型在工业物体上泛化能力差。本文首次验证了此类模型在常见工业物体上的表现不佳。为此,我们提出一种无需训练的开放词汇3D感知流程:不使用预训练模型生成实例提议,而是通过合并预计算的超像素,依据其语义特征生成掩码。随后,我们在代表性3D工业车间场景中评估了领域自适应的VLFM「IndustrialCLIP」的开放词汇查询性能。定性结果表明,该方法成功实现了工业物体的分割。

原文摘要 · Abstract (English)

Autonomous vision applications in production, intralogistics, or manufacturing environments require perception capabilities beyond a small, fixed set of classes. Recent open-vocabulary methods, leveraging 2D Vision-Language Foundation Models (VLFMs), target this task but often rely on class-agnostic segmentation models pre-trained on non-industrial datasets (e.g., household scenes). In this work, we first demonstrate that such models fail to generalize, performing poorly on common industrial objects. Therefore, we propose a training-free, open-vocabulary 3D perception pipeline that overcomes this limitation. Instead of using a pre-trained model to generate instance proposals, our method simply generates masks by merging pre-computed superpoints based on their semantic features. Following, we evaluate the domain-adapted VLFM "IndustrialCLIP" on a representative 3D industrial workshop scene for open-vocabulary querying. Our qualitative results demonstrate successful segmentation of industrial objects.

3D感知开放词汇工业视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。